서브메뉴
검색
Effectively Learning From Data and Generating Data in Differentially Private Machine Learning
Effectively Learning From Data and Generating Data in Differentially Private Machine Learning
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211152825
- ISBN
- 9798384463894
- DDC
- 621.3
- 저자명
- Tang, Xinyu.
- 서명/저자
- Effectively Learning From Data and Generating Data in Differentially Private Machine Learning
- 발행사항
- [Sl] : Princeton University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 190 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-04, Section: B.
- 주기사항
- Advisor: Mittal, Prateek.
- 학위논문주기
- Thesis (Ph.D.)--Princeton University, 2024.
- 초록/해제
- 요약Machine learning models are susceptible to a range of attacks that exploit data leakage from trained models. Differential Privacy (DP) is the gold standard for quantifying privacy risks and providing provable guarantees against attacks. However, training machine learning models with differential privacy often incurs a significant utility drop.In this dissertation, we investigate how to effectively learn from data and generate data in differentially private machine learning. To effectively learn from data in a privacy-preserving way, it is important to identify what kind of prior information we can leverage. Firstly, we study the label-DP set-up, where the feature information is public and the label is private. We investigate how to improve the model utility under label-DP by leveraging public features to add less noise and reduce the effect of noise. Secondly, we study how to leverage the synthetic images to improve differentially private image classification. While such synthetic images are generated without access to real-world images and are only marginally helpful in non-private training, we find that these synthetic images can provide a better prior for differentially private image classification. We further study how to maximize the use of such synthetic priors to further unlock their full potential to improve private training. Thirdly, we study the privatization of zeroth-order optimization which has been shown to achieve competitive performance as SGD in fine-tuning large language models and propose DP-ZO. Our key insight is that in zeroth-order optimization the only information derived from private data is a scalar. Therefore we only need to privatize this scalar quantity. This is privacy friendly as we only need to add noise to a scalar instead of high-dimension gradients. Fourthly, for deferentially private synthetic data generation, we study privately generating the data with API access only to large language models without fine-tuning. Our proposed method can provide the privacy-protection for in-context learning in large language model with unlimited queries.In summary, this dissertation investigates how to effectively learn from data and generate data in differentially private machine learning and provides directions in designing privacy-preserving machine learning models in practice.
- 일반주제명
- Computer engineering
- 일반주제명
- Electrical engineering
- 일반주제명
- Computer science
- 키워드
- Machine learning
- 키워드
- Privacy risks
- 키워드
- Synthetic data
- 기타저자
- Princeton University Electrical and Computer Engineering
- 기본자료저록
- Dissertations Abstracts International. 86-04B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017164050
■00520250211152825
■006m o d
■007cr#unu||||||||
■020 ▼a9798384463894
■035 ▼a(MiAaPQ)AAI31560045
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a621.3
■1001 ▼aTang, Xinyu.
■24510▼aEffectively Learning From Data and Generating Data in Differentially Private Machine Learning
■260 ▼a[Sl]▼bPrinceton University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a190 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-04, Section: B.
■500 ▼aAdvisor: Mittal, Prateek.
■5021 ▼aThesis (Ph.D.)--Princeton University, 2024.
■520 ▼aMachine learning models are susceptible to a range of attacks that exploit data leakage from trained models. Differential Privacy (DP) is the gold standard for quantifying privacy risks and providing provable guarantees against attacks. However, training machine learning models with differential privacy often incurs a significant utility drop.In this dissertation, we investigate how to effectively learn from data and generate data in differentially private machine learning. To effectively learn from data in a privacy-preserving way, it is important to identify what kind of prior information we can leverage. Firstly, we study the label-DP set-up, where the feature information is public and the label is private. We investigate how to improve the model utility under label-DP by leveraging public features to add less noise and reduce the effect of noise. Secondly, we study how to leverage the synthetic images to improve differentially private image classification. While such synthetic images are generated without access to real-world images and are only marginally helpful in non-private training, we find that these synthetic images can provide a better prior for differentially private image classification. We further study how to maximize the use of such synthetic priors to further unlock their full potential to improve private training. Thirdly, we study the privatization of zeroth-order optimization which has been shown to achieve competitive performance as SGD in fine-tuning large language models and propose DP-ZO. Our key insight is that in zeroth-order optimization the only information derived from private data is a scalar. Therefore we only need to privatize this scalar quantity. This is privacy friendly as we only need to add noise to a scalar instead of high-dimension gradients. Fourthly, for deferentially private synthetic data generation, we study privately generating the data with API access only to large language models without fine-tuning. Our proposed method can provide the privacy-protection for in-context learning in large language model with unlimited queries.In summary, this dissertation investigates how to effectively learn from data and generate data in differentially private machine learning and provides directions in designing privacy-preserving machine learning models in practice.
■590 ▼aSchool code: 0181.
■650 4▼aComputer engineering
■650 4▼aElectrical engineering
■650 4▼aComputer science
■653 ▼aDifferential Privacy
■653 ▼aMachine learning
■653 ▼aPrivacy risks
■653 ▼aSynthetic data
■653 ▼aLarge language models
■690 ▼a0464
■690 ▼a0544
■690 ▼a0984
■71020▼aPrinceton University▼bElectrical and Computer Engineering.
■7730 ▼tDissertations Abstracts International▼g86-04B.
■790 ▼a0181
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164050▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


