서브메뉴
검색
Robust Task Specification for Learning Systems
Robust Task Specification for Learning Systems
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211151458
- ISBN
- 9798384450481
- DDC
- 004
- 저자명
- Toyer, Sam.
- 서명/저자
- Robust Task Specification for Learning Systems
- 발행사항
- [Sl] : University of California, Berkeley, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 206 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-03, Section: B.
- 주기사항
- Advisor: Russell, Stuart.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Berkeley, 2024.
- 초록/해제
- 요약This dissertation considers how to evaluate and improve the robustness of AI systems in situations that are systematically different to those encountered during training. Specifically, we focus on test-time robustness for two particular ways of specifying tasks, and two specific forms of generalization. The first part of this dissertation focuses on learning tasks from demonstrations with imitation, while the second focuses on specifying tasks for large language models using natural language instructions.In the first part, we specifically consider the combinatorial and in-distribution generalization of imitation learning. Our first contribution is a benchmark for how well learned policies can generalize along various axes. The benchmark allows us to manipulate these axes independently to determine invariances and equivariances the policy has. Using this benchmark, we show that some basic computer vision techniques (augmentation, egocentric views) improve imitative generalization, but more sophisticated representation learning techniques do not.In the second part, we consider instruction-following language models and adversarial robustness, where a user is actively trying to provoke errors from the model. Here we contribute a large dataset of prompt injection attacks obtained from an online game, which we distill into a benchmark for language model robustness. We also consider a second type of adversarial attack called a jailbreak, and show that existing evaluations are insufficient to gauge the actual misuse potential of jailbreaking techniques. Thus we propose a new benchmark that identifies effective jailbreaks while correctly disregarding ineffective ones.This dissertation proposes several evaluations for challenging problems where existing algorithms fail: imitation learning algorithms struggle to generalize when only few demonstrations are available, and representation learning is not an easy fix. Likewise, the safeguards around large language models are easy for an adversary to subvert. These negative results point toward ways that AI systems could be improved to be more robust in unexpected circumstances; we describe these opportunities for future work in Chapter 6.
- 일반주제명
- Computer science
- 일반주제명
- Computer engineering
- 키워드
- Natural language
- 키워드
- Egocentric views
- 키워드
- Algorithms
- 기타저자
- University of California, Berkeley Computer Science
- 기본자료저록
- Dissertations Abstracts International. 86-03B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017161886
■00520250211151458
■006m o d
■007cr#unu||||||||
■020 ▼a9798384450481
■035 ▼a(MiAaPQ)AAI31297380
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aToyer, Sam.
■24510▼aRobust Task Specification for Learning Systems
■260 ▼a[Sl]▼bUniversity of California, Berkeley▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a206 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-03, Section: B.
■500 ▼aAdvisor: Russell, Stuart.
■5021 ▼aThesis (Ph.D.)--University of California, Berkeley, 2024.
■520 ▼aThis dissertation considers how to evaluate and improve the robustness of AI systems in situations that are systematically different to those encountered during training. Specifically, we focus on test-time robustness for two particular ways of specifying tasks, and two specific forms of generalization. The first part of this dissertation focuses on learning tasks from demonstrations with imitation, while the second focuses on specifying tasks for large language models using natural language instructions.In the first part, we specifically consider the combinatorial and in-distribution generalization of imitation learning. Our first contribution is a benchmark for how well learned policies can generalize along various axes. The benchmark allows us to manipulate these axes independently to determine invariances and equivariances the policy has. Using this benchmark, we show that some basic computer vision techniques (augmentation, egocentric views) improve imitative generalization, but more sophisticated representation learning techniques do not.In the second part, we consider instruction-following language models and adversarial robustness, where a user is actively trying to provoke errors from the model. Here we contribute a large dataset of prompt injection attacks obtained from an online game, which we distill into a benchmark for language model robustness. We also consider a second type of adversarial attack called a jailbreak, and show that existing evaluations are insufficient to gauge the actual misuse potential of jailbreaking techniques. Thus we propose a new benchmark that identifies effective jailbreaks while correctly disregarding ineffective ones.This dissertation proposes several evaluations for challenging problems where existing algorithms fail: imitation learning algorithms struggle to generalize when only few demonstrations are available, and representation learning is not an easy fix. Likewise, the safeguards around large language models are easy for an adversary to subvert. These negative results point toward ways that AI systems could be improved to be more robust in unexpected circumstances; we describe these opportunities for future work in Chapter 6.
■590 ▼aSchool code: 0028.
■650 4▼aComputer science
■650 4▼aComputer engineering
■653 ▼aNatural language
■653 ▼aJailbreaking techniques
■653 ▼aEgocentric views
■653 ▼aAlgorithms
■653 ▼aComputer vision techniques
■690 ▼a0800
■690 ▼a0984
■690 ▼a0464
■71020▼aUniversity of California, Berkeley▼bComputer Science.
■7730 ▼tDissertations Abstracts International▼g86-03B.
■790 ▼a0028
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17161886▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


