본문

서브메뉴

Scalable Auditing for AI Safety
Scalable Auditing for AI Safety
Scalable Auditing for AI Safety

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202103551
ISBN  
9798288866616
DDC  
621.3
저자명  
Jones, Erik.
서명/저자  
Scalable Auditing for AI Safety
발행사항  
[Sl] : University of California, Berkeley, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
197 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-01, Section: B.
주기사항  
Advisor: Steinhardt, Jacob;Dragan, Anca D.
학위논문주기  
Thesis (Ph.D.)--University of California, Berkeley, 2025.
초록/해제  
요약Despite their promise, contemporary AI systems pose safety risks; for example, these systems could be misused by adversaries to conduct malicious tasks, or exhibit behavior that is misaligned with developer intent. However, as both capabilities and deployments scale, effective audits for such risks are becoming increasingly intractable for humans alone to conduct. This is because the risk profile of these systems is increasingly broad: systems may only exhibit certain failures rarely; some failures may be challenging to anticipate a priori; and some failures only emerge in broader contexts. In this thesis, we develop evaluation systems to conduct scalable audits for AI safety. We first aim to develop systems to elicit rare failures-failures that occur sufficiently infrequently that humans might not find them with manual testing. Specifically, we present ARCA, a method that casts auditing for rare failures as a discrete optimization problem over prompts and outputs, which we solve with a novel optimizer. We next develop systems to uncover unexpected failure modes-failures that humans would not have anticipated and tested for beforehand. Specifically, we present MultiMon and TED: two evaluation systems that uncover unforeseen failure modes by studying the relationship between classes of system outputs, rather than assessing the veracity of outputs directly. We finally explore auditing for failures given broader context, and introduce a class of attacks that combines individually-safe systems to produce harmful outputs.
일반주제명  
Computer engineering
키워드  
Scalable audits
키워드  
AI safety
기타저자  
University of California, Berkeley Computer Science
기본자료저록  
Dissertations Abstracts International. 87-01B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017357720
■00520260202103551
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798288866616
■035    ▼a(MiAaPQ)AAI32041820
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a621.3
■1001  ▼aJones,  Erik.
■24510▼aScalable  Auditing  for  AI  Safety
■260    ▼a[Sl]▼bUniversity  of  California,  Berkeley▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a197  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-01,  Section:  B.
■500    ▼aAdvisor:  Steinhardt,  Jacob;Dragan,  Anca  D.
■5021  ▼aThesis  (Ph.D.)--University  of  California,  Berkeley,  2025.
■520    ▼aDespite  their  promise,  contemporary  AI  systems  pose  safety  risks;  for  example,  these  systems  could  be  misused  by  adversaries  to  conduct  malicious  tasks,  or  exhibit  behavior  that  is  misaligned  with  developer  intent.  However,  as  both  capabilities  and  deployments  scale,  effective  audits  for  such  risks  are  becoming  increasingly  intractable  for  humans  alone  to  conduct.  This  is  because  the  risk  profile  of  these  systems  is  increasingly  broad:  systems  may  only  exhibit  certain  failures  rarely;  some  failures  may  be  challenging  to  anticipate  a  priori;  and  some  failures  only  emerge  in  broader  contexts.  In  this  thesis,  we  develop  evaluation  systems  to  conduct  scalable  audits  for  AI  safety.  We  first  aim  to  develop  systems  to  elicit  rare  failures-failures  that  occur  sufficiently  infrequently  that  humans  might  not  find  them  with  manual  testing.  Specifically,  we  present  ARCA,  a  method  that  casts  auditing  for  rare  failures  as  a  discrete  optimization  problem  over  prompts  and  outputs,  which  we  solve  with  a  novel  optimizer.  We  next  develop  systems  to  uncover  unexpected  failure  modes-failures  that  humans  would  not  have  anticipated  and  tested  for  beforehand.  Specifically,  we  present  MultiMon  and  TED:  two  evaluation  systems  that  uncover  unforeseen  failure  modes  by  studying  the  relationship  between  classes  of  system  outputs,  rather  than  assessing  the  veracity  of  outputs  directly.  We  finally  explore  auditing  for  failures  given  broader  context,  and  introduce  a  class  of  attacks  that  combines  individually-safe  systems  to  produce  harmful  outputs.
■590    ▼aSchool  code:  0028.
■650  4▼aComputer  engineering
■653    ▼aScalable  audits
■653    ▼aAI  safety
■690    ▼a0800
■690    ▼a0464
■71020▼aUniversity  of  California,  Berkeley▼bComputer  Science.
■7730  ▼tDissertations  Abstracts  International▼g87-01B.
■790    ▼a0028
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357720▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF18949 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.