서브메뉴
검색
Adversarial Robustness for Estimation and Alignment
Adversarial Robustness for Estimation and Alignment
Detailed Information
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211151333
- ISBN
- 9798382830964
- DDC
- 310
- 저자명
- Chao, Patrick.
- 서명/저자
- Adversarial Robustness for Estimation and Alignment
- 발행사항
- [Sl] : University of Pennsylvania, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 216 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 85-12, Section: B.
- 주기사항
- Advisor: Dobriban, Edgar.
- 학위논문주기
- Thesis (Ph.D.)--University of Pennsylvania, 2024.
- 초록/해제
- 요약As machine learning models are deployed in a multitude of settings with increasing levels of influence and competency, there is growing interest in ensuring these models are robust and align with human intentions. To this end, we analyze robust models and adversarial inputs in a variety of settings. We explore statistical estimation under the adversarial setting of Wasserstein distribution shifts, where every data point may undergo a bounded perturbation. We analyze several statistical problems, including location estimation, linear regression, and non-parametric density estimation. Furthermore, we evaluate alignment in modern foundation models, and propose automated methods to construct adversarial inputs. We develop black-box automated algorithms to generate adversarial prompts for text-to-image models and jailbreaks for language models. Lastly, we introduce a benchmark, JailbreakBench, for reproducible jailbreak evaluation.
- 일반주제명
- Statistics
- 일반주제명
- Information technology
- 키워드
- Jailbreaking
- 키워드
- Red teaming
- 기타저자
- University of Pennsylvania Statistics and Data Science
- 기본자료저록
- Dissertations Abstracts International. 85-12B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017161278
■00520250211151333
■006m o d
■007cr#unu||||||||
■020 ▼a9798382830964
■035 ▼a(MiAaPQ)AAI31241524
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a310
■1001 ▼aChao, Patrick.
■24510▼aAdversarial Robustness for Estimation and Alignment
■260 ▼a[Sl]▼bUniversity of Pennsylvania▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a216 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 85-12, Section: B.
■500 ▼aAdvisor: Dobriban, Edgar.
■5021 ▼aThesis (Ph.D.)--University of Pennsylvania, 2024.
■520 ▼aAs machine learning models are deployed in a multitude of settings with increasing levels of influence and competency, there is growing interest in ensuring these models are robust and align with human intentions. To this end, we analyze robust models and adversarial inputs in a variety of settings. We explore statistical estimation under the adversarial setting of Wasserstein distribution shifts, where every data point may undergo a bounded perturbation. We analyze several statistical problems, including location estimation, linear regression, and non-parametric density estimation. Furthermore, we evaluate alignment in modern foundation models, and propose automated methods to construct adversarial inputs. We develop black-box automated algorithms to generate adversarial prompts for text-to-image models and jailbreaks for language models. Lastly, we introduce a benchmark, JailbreakBench, for reproducible jailbreak evaluation.
■590 ▼aSchool code: 0175.
■650 4▼aStatistics
■650 4▼aInformation technology
■653 ▼aAdversarial prompts
■653 ▼aAdversarial robustness
■653 ▼aDistribution shifts
■653 ▼aJailbreaking
■653 ▼aMinimax estimation
■653 ▼aRed teaming
■690 ▼a0800
■690 ▼a0463
■690 ▼a0489
■71020▼aUniversity of Pennsylvania▼bStatistics and Data Science.
■7730 ▼tDissertations Abstracts International▼g85-12B.
■790 ▼a0175
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17161278▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.
Preview
Export
ChatGPT Discussion
AI Recommended Related Books
detalle info
- Reserva
- No existe
- Mi carpeta
- Primera solicitud
- Non-Book Loan Application
- Nighttime Book Loan Application
Available after logging in.


