본문

서브메뉴

Unlocking Trustworthy Machine Learning With Sparsity
Unlocking Trustworthy Machine Learning With Sparsity
Unlocking Trustworthy Machine Learning With Sparsity

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211152821
ISBN  
9798384463917
DDC  
621.3
저자명  
Panda, Ashwinee.
서명/저자  
Unlocking Trustworthy Machine Learning With Sparsity
발행사항  
[Sl] : Princeton University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
207 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-04, Section: B.
주기사항  
Advisor: Mittal, Prateek.
학위논문주기  
Thesis (Ph.D.)--Princeton University, 2024.
초록/해제  
요약Modern artificial intelligence has come to be defined by the use of large neural networks trained on web-scale datasets. By virtue of their size, these training datasets have become impossible to filter. We show that these models can regurgitate private data verbatim, and pick up harmful or toxic behaviors, even when private or harmful data is scarcely represented in the training data. Understanding why and how models learn undesirable behaviors, such as memorizing private data or toxic behavior, has become a central question in the community of AI safety, as a first step towards mitigating the memorization of private data and preventing models from generating harmful responses. Our underlying insight is that some parameters in an overparametrized model are naturally sparsely updated, meaning that those parameters are only used to fit some out-of-distribution data. This enables attackers to craft data or gradient updates that target those sparsely updated parameters, changing the model's behavior on a small set of inputs. In the same vein, we propose sparsity-based defenses to limit the update surface of a model to the parameters that are most important for the bulk of training. A foundational concept in computer security is limiting the attack surface, and we implement this in machine learning via sparse training. We apply our methodology of sparsity in attacks and defenses to prevent models from learning harmful behavior or memorizing private data.
일반주제명  
Computer engineering
일반주제명  
Electrical engineering
키워드  
Neural networks
키워드  
Toxic behaviors
키워드  
Computer security
키워드  
Machine learning
기타저자  
Princeton University Electrical and Computer Engineering
기본자료저록  
Dissertations Abstracts International. 86-04B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017164017
■00520250211152821
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798384463917
■035    ▼a(MiAaPQ)AAI31559437
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a621.3
■1001  ▼aPanda,  Ashwinee.
■24510▼aUnlocking  Trustworthy  Machine  Learning  With  Sparsity
■260    ▼a[Sl]▼bPrinceton  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a207  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-04,  Section:  B.
■500    ▼aAdvisor:  Mittal,  Prateek.
■5021  ▼aThesis  (Ph.D.)--Princeton  University,  2024.
■520    ▼aModern  artificial  intelligence  has  come  to  be  defined  by  the  use  of  large  neural  networks  trained  on  web-scale  datasets.  By  virtue  of  their  size,  these  training  datasets  have  become  impossible  to  filter.  We  show  that  these  models  can  regurgitate  private  data  verbatim,  and  pick  up  harmful  or  toxic  behaviors,  even  when  private  or  harmful  data  is  scarcely  represented  in  the  training  data.  Understanding  why  and  how  models  learn  undesirable  behaviors,  such  as  memorizing  private  data  or  toxic  behavior,  has  become  a  central  question  in  the  community  of  AI  safety,  as  a  first  step  towards  mitigating  the  memorization  of  private  data  and  preventing  models  from  generating  harmful  responses.  Our  underlying  insight  is  that  some  parameters  in  an  overparametrized  model  are  naturally  sparsely  updated,  meaning  that  those  parameters  are  only  used  to  fit  some  out-of-distribution  data.  This  enables  attackers  to  craft  data  or  gradient  updates  that  target  those  sparsely  updated  parameters,  changing  the  model's  behavior  on  a  small  set  of  inputs.  In  the  same  vein,  we  propose  sparsity-based  defenses  to  limit  the  update  surface  of  a  model  to  the  parameters  that  are  most  important  for  the  bulk  of  training.  A  foundational  concept  in  computer  security  is  limiting  the  attack  surface,  and  we  implement  this  in  machine  learning  via  sparse  training.  We  apply  our  methodology  of  sparsity  in  attacks  and  defenses  to  prevent  models  from  learning  harmful  behavior  or  memorizing  private  data.
■590    ▼aSchool  code:  0181.
■650  4▼aComputer  engineering
■650  4▼aElectrical  engineering
■653    ▼aNeural  networks
■653    ▼aToxic  behaviors
■653    ▼aComputer  security
■653    ▼aMachine  learning
■690    ▼a0800
■690    ▼a0544
■690    ▼a0464
■71020▼aPrinceton  University▼bElectrical  and  Computer  Engineering.
■7730  ▼tDissertations  Abstracts  International▼g86-04B.
■790    ▼a0181
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164017▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF11831 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.