서브메뉴
검색
Principled Machine Learning Under Constraints on Data Access, Quality, and Computations
Principled Machine Learning Under Constraints on Data Access, Quality, and Computations
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260311091548.5
- ISBN
- 9798270232450
- DDC
- 005
- 저자명
- Das, Rudrajit
- 서명/저자
- Principled Machine Learning Under Constraints on Data Access, Quality, and Computations / Rudrajit Das
- 발행사항
- [Sl] : The University of Texas at Austin, 2025
- 형태사항
- 1 electronic resource (351 pages)
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-06, Section: B.
- 주기사항
- Advisors: Sanghavi, Sujay; Dhillon, Inderjit S. Committee members: Liu, Qiang; Kale, Satyen.
- 학위논문주기
- - Ph.D. : The University of Texas at Austin, 2025.
- 초록/해제
- 요약We consider two prevalent data-centric constraints in modern machine learning: (a) restricted data access with potential computational constraints, and (b) poor data quality. Our goal is to provide theoretically sound algorithms/practices for such settings. Under (a), we focus on federated learning (FL) where data is stored locally on decentralized clients, each with individual computational constraints, and on differentially private training where data access is impaired due to the privacy-preservation requirement. Specifically, we propose an accelerated FL algorithm attaining the best known complexity for smooth non-convex functions under arbitrary client heterogeneity and compressed communication. We also provide a theoretically justified recommendation for setting the clip norm in differentially private stochastic gradient descent (DP-SGD) and derive new convergence results for DP-SGD with heavy-tailed gradients. We validate the effectiveness of our methods via extensive experimentation. Under (b), we consider the problem of learning with noisy labels in this dissertation. Specifically, we focus on the idea of retraining a model with its own hard predictions (1/0 labels) or soft predictions (raw unrounded scores) on the same training set on which it is initially trained. Surprisingly, this simple idea improves the model's performance, even though no extra information is obtained by retraining. We theoretically characterize this surprising phenomenon for linear models; to our knowledge, our results are the first of their kind. Empirically, we show the efficacy of selective retraining in improving training with local label differential privacy, where the goal is to safeguard the privacy of only the labels by injecting label noise.
- 언어주기
- English
- 일반주제명
- Computer science
- 일반주제명
- Information technology
- 일반주제명
- Information science
- 키워드
- Machine learning
- 키워드
- Data access
- 키워드
- Data quality
- 기타저자
- The University of Texas at Austin Computer Science
- 기본자료저록
- Dissertations Abstracts International. 87-06B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260311s2025 us eng d■001000017361237
■00520260311091548.5
■006m o d
■007cr|nu||||||||
■020 ▼a9798270232450
■040 ▼aMiAaPQD▼beng▼cMiAaPQD▼erda
■082 ▼a005
■1001 ▼aDas, Rudrajit▼eauthor.
■24510▼aPrincipled Machine Learning Under Constraints on Data Access, Quality, and Computations ▼cRudrajit Das
■260 ▼a[Sl]▼bThe University of Texas at Austin▼c2025
■264 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a1 electronic resource (351 pages)
■336 ▼atext▼btxt▼2rdacontent
■337 ▼acomputer▼bc▼2rdamedia
■338 ▼aonline resource▼bcr▼2rdacarrier
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-06, Section: B.
■500 ▼aAdvisors: Sanghavi, Sujay; Dhillon, Inderjit S. Committee members: Liu, Qiang; Kale, Satyen.
■5021 ▼bPh.D.▼cThe University of Texas at Austin▼d2025.
■520 ▼aWe consider two prevalent data-centric constraints in modern machine learning: (a) restricted data access with potential computational constraints, and (b) poor data quality. Our goal is to provide theoretically sound algorithms/practices for such settings. Under (a), we focus on federated learning (FL) where data is stored locally on decentralized clients, each with individual computational constraints, and on differentially private training where data access is impaired due to the privacy-preservation requirement. Specifically, we propose an accelerated FL algorithm attaining the best known complexity for smooth non-convex functions under arbitrary client heterogeneity and compressed communication. We also provide a theoretically justified recommendation for setting the clip norm in differentially private stochastic gradient descent (DP-SGD) and derive new convergence results for DP-SGD with heavy-tailed gradients. We validate the effectiveness of our methods via extensive experimentation. Under (b), we consider the problem of learning with noisy labels in this dissertation. Specifically, we focus on the idea of retraining a model with its own hard predictions (1/0 labels) or soft predictions (raw unrounded scores) on the same training set on which it is initially trained. Surprisingly, this simple idea improves the model's performance, even though no extra information is obtained by retraining. We theoretically characterize this surprising phenomenon for linear models; to our knowledge, our results are the first of their kind. Empirically, we show the efficacy of selective retraining in improving training with local label differential privacy, where the goal is to safeguard the privacy of only the labels by injecting label noise.
■546 ▼aEnglish
■590 ▼aSchool code: 0227
■650 4▼aComputer science
■650 4▼aInformation technology
■650 4▼aInformation science
■653 ▼aMachine learning
■653 ▼aFederated learning
■653 ▼aData access
■653 ▼aData quality
■7102 ▼aThe University of Texas at Austin▼bComputer Science.▼edegree granting institution.
■7201 ▼aSanghavi, Sujay▼edegree supervisor.
■7201 ▼aDhillon, Inderjit S.▼edegree supervisor.
■7730 ▼tDissertations Abstracts International▼g87-06B.
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17361237▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


