본문

서브메뉴

Robust Task Specification for Learning Systems
Robust Task Specification for Learning Systems
Robust Task Specification for Learning Systems

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211151458
ISBN  
9798384450481
DDC  
004
저자명  
Toyer, Sam.
서명/저자  
Robust Task Specification for Learning Systems
발행사항  
[Sl] : University of California, Berkeley, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
206 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-03, Section: B.
주기사항  
Advisor: Russell, Stuart.
학위논문주기  
Thesis (Ph.D.)--University of California, Berkeley, 2024.
초록/해제  
요약This dissertation considers how to evaluate and improve the robustness of AI systems in situations that are systematically different to those encountered during training. Specifically, we focus on test-time robustness for two particular ways of specifying tasks, and two specific forms of generalization. The first part of this dissertation focuses on learning tasks from demonstrations with imitation, while the second focuses on specifying tasks for large language models using natural language instructions.In the first part, we specifically consider the combinatorial and in-distribution generalization of imitation learning. Our first contribution is a benchmark for how well learned policies can generalize along various axes. The benchmark allows us to manipulate these axes independently to determine invariances and equivariances the policy has. Using this benchmark, we show that some basic computer vision techniques (augmentation, egocentric views) improve imitative generalization, but more sophisticated representation learning techniques do not.In the second part, we consider instruction-following language models and adversarial robustness, where a user is actively trying to provoke errors from the model. Here we contribute a large dataset of prompt injection attacks obtained from an online game, which we distill into a benchmark for language model robustness. We also consider a second type of adversarial attack called a jailbreak, and show that existing evaluations are insufficient to gauge the actual misuse potential of jailbreaking techniques. Thus we propose a new benchmark that identifies effective jailbreaks while correctly disregarding ineffective ones.This dissertation proposes several evaluations for challenging problems where existing algorithms fail: imitation learning algorithms struggle to generalize when only few demonstrations are available, and representation learning is not an easy fix. Likewise, the safeguards around large language models are easy for an adversary to subvert. These negative results point toward ways that AI systems could be improved to be more robust in unexpected circumstances; we describe these opportunities for future work in Chapter 6.
일반주제명  
Computer science
일반주제명  
Computer engineering
키워드  
Natural language
키워드  
Jailbreaking techniques
키워드  
Egocentric views
키워드  
Algorithms
키워드  
Computer vision techniques
기타저자  
University of California, Berkeley Computer Science
기본자료저록  
Dissertations Abstracts International. 86-03B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017161886
■00520250211151458
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798384450481
■035    ▼a(MiAaPQ)AAI31297380
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aToyer,  Sam.
■24510▼aRobust  Task  Specification  for  Learning  Systems
■260    ▼a[Sl]▼bUniversity  of  California,  Berkeley▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a206  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-03,  Section:  B.
■500    ▼aAdvisor:  Russell,  Stuart.
■5021  ▼aThesis  (Ph.D.)--University  of  California,  Berkeley,  2024.
■520    ▼aThis  dissertation  considers  how  to  evaluate  and  improve  the  robustness  of  AI  systems  in  situations  that  are  systematically  different  to  those  encountered  during  training.  Specifically,  we  focus  on  test-time  robustness  for  two  particular  ways  of  specifying  tasks,  and  two  specific  forms  of  generalization.  The  first  part  of  this  dissertation  focuses  on  learning  tasks  from  demonstrations  with  imitation,  while  the  second  focuses  on  specifying  tasks  for  large  language  models  using  natural  language  instructions.In  the  first  part,  we  specifically  consider  the  combinatorial  and  in-distribution  generalization  of  imitation  learning.  Our  first  contribution  is  a  benchmark  for  how  well  learned  policies  can  generalize  along  various  axes.  The  benchmark  allows  us  to  manipulate  these  axes  independently  to  determine  invariances  and  equivariances  the  policy  has.  Using  this  benchmark,  we  show  that  some  basic  computer  vision  techniques  (augmentation,  egocentric  views)  improve  imitative  generalization,  but  more  sophisticated  representation  learning  techniques  do  not.In  the  second  part,  we  consider  instruction-following  language  models  and  adversarial  robustness,  where  a  user  is  actively  trying  to  provoke  errors  from  the  model.  Here  we  contribute  a  large  dataset  of  prompt  injection  attacks  obtained  from  an  online  game,  which  we  distill  into  a  benchmark  for  language  model  robustness.  We  also  consider  a  second  type  of  adversarial  attack  called  a  jailbreak,  and  show  that  existing  evaluations  are  insufficient  to  gauge  the  actual  misuse  potential  of  jailbreaking  techniques.  Thus  we  propose  a  new  benchmark  that  identifies  effective  jailbreaks  while  correctly  disregarding  ineffective  ones.This  dissertation  proposes  several  evaluations  for  challenging  problems  where  existing  algorithms  fail:  imitation  learning  algorithms  struggle  to  generalize  when  only  few  demonstrations  are  available,  and  representation  learning  is  not  an  easy  fix.  Likewise,  the  safeguards  around  large  language  models  are  easy  for  an  adversary  to  subvert.  These  negative  results  point  toward  ways  that  AI  systems  could  be  improved  to  be  more  robust  in  unexpected  circumstances;  we  describe  these  opportunities  for  future  work  in  Chapter  6.
■590    ▼aSchool  code:  0028.
■650  4▼aComputer  science
■650  4▼aComputer  engineering
■653    ▼aNatural  language
■653    ▼aJailbreaking  techniques
■653    ▼aEgocentric  views
■653    ▼aAlgorithms
■653    ▼aComputer  vision  techniques
■690    ▼a0800
■690    ▼a0984
■690    ▼a0464
■71020▼aUniversity  of  California,  Berkeley▼bComputer  Science.
■7730  ▼tDissertations  Abstracts  International▼g86-03B.
■790    ▼a0028
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17161886▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF10876 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.