서브메뉴
검색
Unlocking Trustworthy Machine Learning With Sparsity
Unlocking Trustworthy Machine Learning With Sparsity
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211152821
- ISBN
- 9798384463917
- DDC
- 621.3
- 저자명
- Panda, Ashwinee.
- 서명/저자
- Unlocking Trustworthy Machine Learning With Sparsity
- 발행사항
- [Sl] : Princeton University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 207 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-04, Section: B.
- 주기사항
- Advisor: Mittal, Prateek.
- 학위논문주기
- Thesis (Ph.D.)--Princeton University, 2024.
- 초록/해제
- 요약Modern artificial intelligence has come to be defined by the use of large neural networks trained on web-scale datasets. By virtue of their size, these training datasets have become impossible to filter. We show that these models can regurgitate private data verbatim, and pick up harmful or toxic behaviors, even when private or harmful data is scarcely represented in the training data. Understanding why and how models learn undesirable behaviors, such as memorizing private data or toxic behavior, has become a central question in the community of AI safety, as a first step towards mitigating the memorization of private data and preventing models from generating harmful responses. Our underlying insight is that some parameters in an overparametrized model are naturally sparsely updated, meaning that those parameters are only used to fit some out-of-distribution data. This enables attackers to craft data or gradient updates that target those sparsely updated parameters, changing the model's behavior on a small set of inputs. In the same vein, we propose sparsity-based defenses to limit the update surface of a model to the parameters that are most important for the bulk of training. A foundational concept in computer security is limiting the attack surface, and we implement this in machine learning via sparse training. We apply our methodology of sparsity in attacks and defenses to prevent models from learning harmful behavior or memorizing private data.
- 일반주제명
- Computer engineering
- 일반주제명
- Electrical engineering
- 키워드
- Neural networks
- 키워드
- Toxic behaviors
- 키워드
- Machine learning
- 기타저자
- Princeton University Electrical and Computer Engineering
- 기본자료저록
- Dissertations Abstracts International. 86-04B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017164017
■00520250211152821
■006m o d
■007cr#unu||||||||
■020 ▼a9798384463917
■035 ▼a(MiAaPQ)AAI31559437
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a621.3
■1001 ▼aPanda, Ashwinee.
■24510▼aUnlocking Trustworthy Machine Learning With Sparsity
■260 ▼a[Sl]▼bPrinceton University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a207 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-04, Section: B.
■500 ▼aAdvisor: Mittal, Prateek.
■5021 ▼aThesis (Ph.D.)--Princeton University, 2024.
■520 ▼aModern artificial intelligence has come to be defined by the use of large neural networks trained on web-scale datasets. By virtue of their size, these training datasets have become impossible to filter. We show that these models can regurgitate private data verbatim, and pick up harmful or toxic behaviors, even when private or harmful data is scarcely represented in the training data. Understanding why and how models learn undesirable behaviors, such as memorizing private data or toxic behavior, has become a central question in the community of AI safety, as a first step towards mitigating the memorization of private data and preventing models from generating harmful responses. Our underlying insight is that some parameters in an overparametrized model are naturally sparsely updated, meaning that those parameters are only used to fit some out-of-distribution data. This enables attackers to craft data or gradient updates that target those sparsely updated parameters, changing the model's behavior on a small set of inputs. In the same vein, we propose sparsity-based defenses to limit the update surface of a model to the parameters that are most important for the bulk of training. A foundational concept in computer security is limiting the attack surface, and we implement this in machine learning via sparse training. We apply our methodology of sparsity in attacks and defenses to prevent models from learning harmful behavior or memorizing private data.
■590 ▼aSchool code: 0181.
■650 4▼aComputer engineering
■650 4▼aElectrical engineering
■653 ▼aNeural networks
■653 ▼aToxic behaviors
■653 ▼aComputer security
■653 ▼aMachine learning
■690 ▼a0800
■690 ▼a0544
■690 ▼a0464
■71020▼aPrinceton University▼bElectrical and Computer Engineering.
■7730 ▼tDissertations Abstracts International▼g86-04B.
■790 ▼a0181
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164017▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


