본문

서브메뉴

Principled Machine Learning Under Constraints on Data Access, Quality, and Computations
Principled Machine Learning Under Constraints on Data Access, Quality, and Computations  /...
Principled Machine Learning Under Constraints on Data Access, Quality, and Computations

Detailed Information

자료유형  
 학위논문 서양
최종처리일시  
20260311091548.5
ISBN  
9798270232450
DDC  
005
저자명  
Das, Rudrajit
서명/저자  
Principled Machine Learning Under Constraints on Data Access, Quality, and Computations / Rudrajit Das
발행사항  
[Sl] : The University of Texas at Austin, 2025
형태사항  
1 electronic resource (351 pages)
주기사항  
Source: Dissertations Abstracts International, Volume: 87-06, Section: B.
주기사항  
Advisors: Sanghavi, Sujay; Dhillon, Inderjit S. Committee members: Liu, Qiang; Kale, Satyen.
학위논문주기  
- Ph.D. : The University of Texas at Austin, 2025.
초록/해제  
요약We consider two prevalent data-centric constraints in modern machine learning: (a) restricted data access with potential computational constraints, and (b) poor data quality. Our goal is to provide theoretically sound algorithms/practices for such settings. Under (a), we focus on federated learning (FL) where data is stored locally on decentralized clients, each with individual computational constraints, and on differentially private training where data access is impaired due to the privacy-preservation requirement. Specifically, we propose an accelerated FL algorithm attaining the best known complexity for smooth non-convex functions under arbitrary client heterogeneity and compressed communication. We also provide a theoretically justified recommendation for setting the clip norm in differentially private stochastic gradient descent (DP-SGD) and derive new convergence results for DP-SGD with heavy-tailed gradients. We validate the effectiveness of our methods via extensive experimentation. Under (b), we consider the problem of learning with noisy labels in this dissertation. Specifically, we focus on the idea of retraining a model with its own hard predictions (1/0 labels) or soft predictions (raw unrounded scores) on the same training set on which it is initially trained. Surprisingly, this simple idea improves the model's performance, even though no extra information is obtained by retraining. We theoretically characterize this surprising phenomenon for linear models; to our knowledge, our results are the first of their kind. Empirically, we show the efficacy of selective retraining in improving training with local label differential privacy, where the goal is to safeguard the privacy of only the labels by injecting label noise.
언어주기  
English
일반주제명  
Computer science
일반주제명  
Information technology
일반주제명  
Information science
키워드  
Machine learning
키워드  
Federated learning
키워드  
Data access
키워드  
Data quality
기타저자  
The University of Texas at Austin Computer Science
기본자료저록  
Dissertations Abstracts International. 87-06B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260311s2025        us                                    eng  d
■001000017361237
■00520260311091548.5
■006m          o    d                
■007cr|nu||||||||
■020    ▼a9798270232450
■040    ▼aMiAaPQD▼beng▼cMiAaPQD▼erda
■082    ▼a005
■1001  ▼aDas,  Rudrajit▼eauthor.
■24510▼aPrincipled  Machine  Learning  Under  Constraints  on  Data  Access,  Quality,  and  Computations  ▼cRudrajit  Das
■260    ▼a[Sl]▼bThe  University  of  Texas  at  Austin▼c2025
■264  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a1  electronic  resource  (351  pages)
■336    ▼atext▼btxt▼2rdacontent
■337    ▼acomputer▼bc▼2rdamedia
■338    ▼aonline  resource▼bcr▼2rdacarrier
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-06,  Section:  B.
■500    ▼aAdvisors:  Sanghavi,  Sujay;  Dhillon,  Inderjit  S.    Committee  members:  Liu,  Qiang;  Kale,  Satyen.
■5021  ▼bPh.D.▼cThe  University  of  Texas  at  Austin▼d2025.
■520    ▼aWe  consider  two  prevalent  data-centric  constraints  in  modern  machine  learning:  (a)  restricted  data  access  with  potential  computational  constraints,  and  (b)  poor  data  quality.  Our  goal  is  to  provide  theoretically  sound  algorithms/practices  for  such  settings.                                                  Under  (a),  we  focus  on  federated  learning  (FL)  where  data  is  stored  locally  on  decentralized  clients,  each  with  individual  computational  constraints,  and  on  differentially  private  training  where  data  access  is  impaired  due  to  the  privacy-preservation  requirement.  Specifically,  we  propose  an  accelerated  FL  algorithm  attaining  the  best  known  complexity  for  smooth  non-convex  functions  under  arbitrary  client  heterogeneity  and  compressed  communication.  We  also  provide  a  theoretically  justified  recommendation  for  setting  the  clip  norm  in  differentially  private  stochastic  gradient  descent  (DP-SGD)  and  derive  new  convergence  results  for  DP-SGD  with  heavy-tailed  gradients.  We  validate  the  effectiveness  of  our  methods  via  extensive  experimentation.                                                  Under  (b),  we  consider  the  problem  of  learning  with  noisy  labels  in  this  dissertation.  Specifically,  we  focus  on  the  idea  of  retraining  a  model  with  its  own  hard  predictions  (1/0  labels)  or  soft  predictions  (raw  unrounded  scores)  on  the  same  training  set  on  which  it  is  initially  trained.  Surprisingly,  this  simple  idea  improves  the  model's  performance,  even  though  no  extra  information  is  obtained  by  retraining.  We  theoretically  characterize  this  surprising  phenomenon  for  linear  models;  to  our  knowledge,  our  results  are  the  first  of  their  kind.  Empirically,  we  show  the  efficacy  of  selective  retraining  in  improving  training  with  local  label  differential  privacy,  where  the  goal  is  to  safeguard  the  privacy  of  only  the  labels  by  injecting  label  noise.
■546    ▼aEnglish
■590    ▼aSchool  code:  0227
■650  4▼aComputer  science
■650  4▼aInformation  technology
■650  4▼aInformation  science
■653    ▼aMachine  learning
■653    ▼aFederated  learning
■653    ▼aData  access
■653    ▼aData  quality
■7102  ▼aThe  University  of  Texas  at  Austin▼bComputer  Science.▼edegree  granting  institution.
■7201  ▼aSanghavi,  Sujay▼edegree  supervisor.
■7201  ▼aDhillon,  Inderjit  S.▼edegree  supervisor.
■7730  ▼tDissertations  Abstracts  International▼g87-06B.
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17361237▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

Preview

Export

ChatGPT Discussion

AI Recommended Related Books


    New Books MORE
    Statistics for the past 3 years. Go to brief

    Подробнее информация.

    • Бронирование
    • не существует
    • моя папка
    • Первый запрос зрения
    • Non-Book Loan Application
    • Nighttime Book Loan Application
    материал
    Reg No. Количество платежных Местоположение статус Ленд информации
    TF18071 전자도서 대출가능 My Folder 부재도서신고 비도서대출신청 야간 도서대출신청

    * Бронирование доступны в заимствований книги. Чтобы сделать предварительный заказ, пожалуйста, нажмите кнопку бронирование

    Books borrowed together with this book

    Related Popular Books

    Available after logging in.