본문

서브메뉴

Effectively Learning From Data and Generating Data in Differentially Private Machine Learning
Effectively Learning From Data and Generating Data in Differentially Private Machine Learn...
Effectively Learning From Data and Generating Data in Differentially Private Machine Learning

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211152825
ISBN  
9798384463894
DDC  
621.3
저자명  
Tang, Xinyu.
서명/저자  
Effectively Learning From Data and Generating Data in Differentially Private Machine Learning
발행사항  
[Sl] : Princeton University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
190 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-04, Section: B.
주기사항  
Advisor: Mittal, Prateek.
학위논문주기  
Thesis (Ph.D.)--Princeton University, 2024.
초록/해제  
요약Machine learning models are susceptible to a range of attacks that exploit data leakage from trained models. Differential Privacy (DP) is the gold standard for quantifying privacy risks and providing provable guarantees against attacks. However, training machine learning models with differential privacy often incurs a significant utility drop.In this dissertation, we investigate how to effectively learn from data and generate data in differentially private machine learning. To effectively learn from data in a privacy-preserving way, it is important to identify what kind of prior information we can leverage. Firstly, we study the label-DP set-up, where the feature information is public and the label is private. We investigate how to improve the model utility under label-DP by leveraging public features to add less noise and reduce the effect of noise. Secondly, we study how to leverage the synthetic images to improve differentially private image classification. While such synthetic images are generated without access to real-world images and are only marginally helpful in non-private training, we find that these synthetic images can provide a better prior for differentially private image classification. We further study how to maximize the use of such synthetic priors to further unlock their full potential to improve private training. Thirdly, we study the privatization of zeroth-order optimization which has been shown to achieve competitive performance as SGD in fine-tuning large language models and propose DP-ZO. Our key insight is that in zeroth-order optimization the only information derived from private data is a scalar. Therefore we only need to privatize this scalar quantity. This is privacy friendly as we only need to add noise to a scalar instead of high-dimension gradients. Fourthly, for deferentially private synthetic data generation, we study privately generating the data with API access only to large language models without fine-tuning. Our proposed method can provide the privacy-protection for in-context learning in large language model with unlimited queries.In summary, this dissertation investigates how to effectively learn from data and generate data in differentially private machine learning and provides directions in designing privacy-preserving machine learning models in practice.
일반주제명  
Computer engineering
일반주제명  
Electrical engineering
일반주제명  
Computer science
키워드  
Differential Privacy
키워드  
Machine learning
키워드  
Privacy risks
키워드  
Synthetic data
키워드  
Large language models
기타저자  
Princeton University Electrical and Computer Engineering
기본자료저록  
Dissertations Abstracts International. 86-04B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017164050
■00520250211152825
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798384463894
■035    ▼a(MiAaPQ)AAI31560045
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a621.3
■1001  ▼aTang,  Xinyu.
■24510▼aEffectively  Learning  From  Data  and  Generating  Data  in  Differentially  Private  Machine  Learning
■260    ▼a[Sl]▼bPrinceton  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a190  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-04,  Section:  B.
■500    ▼aAdvisor:  Mittal,  Prateek.
■5021  ▼aThesis  (Ph.D.)--Princeton  University,  2024.
■520    ▼aMachine  learning  models  are  susceptible  to  a  range  of  attacks  that  exploit  data  leakage  from  trained  models.  Differential  Privacy  (DP)  is  the  gold  standard  for  quantifying  privacy  risks  and  providing  provable  guarantees  against  attacks.  However,  training  machine  learning  models  with  differential  privacy  often  incurs  a  significant  utility  drop.In  this  dissertation,  we  investigate  how  to  effectively  learn  from  data  and  generate  data  in  differentially  private  machine  learning.  To  effectively  learn  from  data  in  a  privacy-preserving  way,  it  is  important  to  identify  what  kind  of  prior  information  we  can  leverage.  Firstly,  we  study  the  label-DP  set-up,  where  the  feature  information  is  public  and  the  label  is  private.  We  investigate  how  to  improve  the  model  utility  under  label-DP  by  leveraging  public  features  to  add  less  noise  and  reduce  the  effect  of  noise.  Secondly,  we  study  how  to  leverage  the  synthetic  images  to  improve  differentially  private  image  classification.  While  such  synthetic  images  are  generated  without  access  to  real-world  images  and  are  only  marginally  helpful  in  non-private  training,  we  find  that  these  synthetic  images  can  provide  a  better  prior  for  differentially  private  image  classification.  We  further  study  how  to  maximize  the  use  of  such  synthetic  priors  to  further  unlock  their  full  potential  to  improve  private  training.  Thirdly,  we  study  the  privatization  of  zeroth-order  optimization  which  has  been  shown  to  achieve  competitive  performance  as  SGD  in  fine-tuning  large  language  models  and  propose  DP-ZO.  Our  key  insight  is  that  in  zeroth-order  optimization  the  only  information  derived  from  private  data  is  a  scalar.  Therefore  we  only  need  to  privatize  this  scalar  quantity.  This  is  privacy  friendly  as  we  only  need  to  add  noise  to  a  scalar  instead  of  high-dimension  gradients.  Fourthly,  for  deferentially  private  synthetic  data  generation,  we  study  privately  generating  the  data  with  API  access  only  to  large  language  models  without  fine-tuning.  Our  proposed  method  can  provide  the  privacy-protection  for  in-context  learning  in  large  language  model  with  unlimited  queries.In  summary,  this  dissertation  investigates  how  to  effectively  learn  from  data  and  generate  data  in  differentially  private  machine  learning  and  provides  directions  in  designing  privacy-preserving  machine  learning  models  in  practice.
■590    ▼aSchool  code:  0181.
■650  4▼aComputer  engineering
■650  4▼aElectrical  engineering
■650  4▼aComputer  science
■653    ▼aDifferential  Privacy
■653    ▼aMachine  learning
■653    ▼aPrivacy  risks
■653    ▼aSynthetic  data
■653    ▼aLarge  language  models
■690    ▼a0464
■690    ▼a0544
■690    ▼a0984
■71020▼aPrinceton  University▼bElectrical  and  Computer  Engineering.
■7730  ▼tDissertations  Abstracts  International▼g86-04B.
■790    ▼a0181
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164050▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF11986 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.