서브메뉴
검색
Scalable Auditing for AI Safety
Scalable Auditing for AI Safety
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103551
- ISBN
- 9798288866616
- DDC
- 621.3
- 저자명
- Jones, Erik.
- 서명/저자
- Scalable Auditing for AI Safety
- 발행사항
- [Sl] : University of California, Berkeley, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 197 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-01, Section: B.
- 주기사항
- Advisor: Steinhardt, Jacob;Dragan, Anca D.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Berkeley, 2025.
- 초록/해제
- 요약Despite their promise, contemporary AI systems pose safety risks; for example, these systems could be misused by adversaries to conduct malicious tasks, or exhibit behavior that is misaligned with developer intent. However, as both capabilities and deployments scale, effective audits for such risks are becoming increasingly intractable for humans alone to conduct. This is because the risk profile of these systems is increasingly broad: systems may only exhibit certain failures rarely; some failures may be challenging to anticipate a priori; and some failures only emerge in broader contexts. In this thesis, we develop evaluation systems to conduct scalable audits for AI safety. We first aim to develop systems to elicit rare failures-failures that occur sufficiently infrequently that humans might not find them with manual testing. Specifically, we present ARCA, a method that casts auditing for rare failures as a discrete optimization problem over prompts and outputs, which we solve with a novel optimizer. We next develop systems to uncover unexpected failure modes-failures that humans would not have anticipated and tested for beforehand. Specifically, we present MultiMon and TED: two evaluation systems that uncover unforeseen failure modes by studying the relationship between classes of system outputs, rather than assessing the veracity of outputs directly. We finally explore auditing for failures given broader context, and introduce a class of attacks that combines individually-safe systems to produce harmful outputs.
- 일반주제명
- Computer engineering
- 키워드
- Scalable audits
- 키워드
- AI safety
- 기타저자
- University of California, Berkeley Computer Science
- 기본자료저록
- Dissertations Abstracts International. 87-01B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017357720
■00520260202103551
■006m o d
■007cr#unu||||||||
■020 ▼a9798288866616
■035 ▼a(MiAaPQ)AAI32041820
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a621.3
■1001 ▼aJones, Erik.
■24510▼aScalable Auditing for AI Safety
■260 ▼a[Sl]▼bUniversity of California, Berkeley▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a197 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-01, Section: B.
■500 ▼aAdvisor: Steinhardt, Jacob;Dragan, Anca D.
■5021 ▼aThesis (Ph.D.)--University of California, Berkeley, 2025.
■520 ▼aDespite their promise, contemporary AI systems pose safety risks; for example, these systems could be misused by adversaries to conduct malicious tasks, or exhibit behavior that is misaligned with developer intent. However, as both capabilities and deployments scale, effective audits for such risks are becoming increasingly intractable for humans alone to conduct. This is because the risk profile of these systems is increasingly broad: systems may only exhibit certain failures rarely; some failures may be challenging to anticipate a priori; and some failures only emerge in broader contexts. In this thesis, we develop evaluation systems to conduct scalable audits for AI safety. We first aim to develop systems to elicit rare failures-failures that occur sufficiently infrequently that humans might not find them with manual testing. Specifically, we present ARCA, a method that casts auditing for rare failures as a discrete optimization problem over prompts and outputs, which we solve with a novel optimizer. We next develop systems to uncover unexpected failure modes-failures that humans would not have anticipated and tested for beforehand. Specifically, we present MultiMon and TED: two evaluation systems that uncover unforeseen failure modes by studying the relationship between classes of system outputs, rather than assessing the veracity of outputs directly. We finally explore auditing for failures given broader context, and introduce a class of attacks that combines individually-safe systems to produce harmful outputs.
■590 ▼aSchool code: 0028.
■650 4▼aComputer engineering
■653 ▼aScalable audits
■653 ▼aAI safety
■690 ▼a0800
■690 ▼a0464
■71020▼aUniversity of California, Berkeley▼bComputer Science.
■7730 ▼tDissertations Abstracts International▼g87-01B.
■790 ▼a0028
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357720▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


