본문

서브메뉴

Towards Data-Efficient Machine Learning Systems
Towards Data-Efficient Machine Learning Systems
Towards Data-Efficient Machine Learning Systems

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211151923
ISBN  
9798346738145
DDC  
004
저자명  
Sachdeva, Noveen.
서명/저자  
Towards Data-Efficient Machine Learning Systems
발행사항  
[Sl] : University of California, San Diego, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
172 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-05, Section: B.
주기사항  
Advisor: McAuley, Julian.
학위논문주기  
Thesis (Ph.D.)--University of California, San Diego, 2024.
초록/해제  
요약The amount of data available to train modern machine learning systems has been increasing rapidly, so much so that we're using, e.g., entirety of the publicly available text data to train state-of-the-art (SoTA) large language models (LLMs), interaction data from billions of users to train SoTA recommender systems, etc. Training of such large machine learning systems on such large datasets entails a high (i) computational runtime, (ii) economical cost, and (iii) carbon footprint; all of which we aim to minimize for different reasons.While a large body of literature develops "model-centric" techniques to better model a given dataset, in this thesis, we develop a "data-centric" viewpoint, where we are interested in techniques that can appropriately summarize a given training dataset, such that models can be trained equally effectively on the data summary vs. training on the much larger original dataset. In addition to being more efficient overall, data-efficient techniques further aim to improve the trained model's quality by stripping away the low-quality and noisy sources of information in the original dataset.More specifically, we develop techniques from two disparate data summarization ideologies: (i) data pruning (a.k.a. coreset construction) techniques that sample the most relevant portions from the dataset using various grounded heuristics, and (ii) data distillation techniques that generate synthetic data-points which summarize the underlying information in the dataset, and are optimized end-to-end using meta-learning. We restrict our scope to training (i) language models on textual datasets, and (ii) recommender systems on user-item interaction datasets.By pushing the frontier of data-efficient training of machine learning systems, we believe our research can effectively contribute to the practical success of such widely-deployed systems, as well as provide a better understanding for the research community to build future work on.
일반주제명  
Computer science
키워드  
Data efficiency
키워드  
Large language models
키워드  
Machine learning
키워드  
Recommender systems
기타저자  
University of California, San Diego Computer Science and Engineering
기본자료저록  
Dissertations Abstracts International. 86-05B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017162134
■00520250211151923
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798346738145
■035    ▼a(MiAaPQ)AAI31295211
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aSachdeva,  Noveen.
■24510▼aTowards  Data-Efficient  Machine  Learning  Systems
■260    ▼a[Sl]▼bUniversity  of  California,  San  Diego▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a172  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-05,  Section:  B.
■500    ▼aAdvisor:  McAuley,  Julian.
■5021  ▼aThesis  (Ph.D.)--University  of  California,  San  Diego,  2024.
■520    ▼aThe  amount  of  data  available  to  train  modern  machine  learning  systems  has  been  increasing  rapidly,  so  much  so  that  we're  using,  e.g.,  entirety  of  the  publicly  available  text  data  to  train  state-of-the-art  (SoTA)  large  language  models  (LLMs),  interaction  data  from  billions  of  users  to  train  SoTA  recommender  systems,  etc.  Training  of  such  large  machine  learning  systems  on  such  large  datasets  entails  a  high  (i)  computational  runtime,  (ii)  economical  cost,  and  (iii)  carbon  footprint;  all  of  which  we  aim  to  minimize  for  different  reasons.While  a  large  body  of  literature  develops  "model-centric"  techniques  to  better  model  a  given  dataset,  in  this  thesis,  we  develop  a  "data-centric"  viewpoint,  where  we  are  interested  in  techniques  that  can  appropriately  summarize  a  given  training  dataset,  such  that  models  can  be  trained  equally  effectively  on  the  data  summary  vs.  training  on  the  much  larger  original  dataset.  In  addition  to  being  more  efficient  overall,  data-efficient  techniques  further  aim  to  improve  the  trained  model's  quality  by  stripping  away  the  low-quality  and  noisy  sources  of  information  in  the  original  dataset.More  specifically,  we  develop  techniques  from  two  disparate  data  summarization  ideologies:  (i)  data  pruning  (a.k.a.  coreset  construction)  techniques  that  sample  the  most  relevant  portions  from  the  dataset  using  various  grounded  heuristics,  and  (ii)  data  distillation  techniques  that  generate  synthetic  data-points  which  summarize  the  underlying  information  in  the  dataset,  and  are  optimized  end-to-end  using  meta-learning.  We  restrict  our  scope  to  training  (i)  language  models  on  textual  datasets,  and  (ii)  recommender  systems  on  user-item  interaction  datasets.By  pushing  the  frontier  of  data-efficient  training  of  machine  learning  systems,  we  believe  our  research  can  effectively  contribute  to  the  practical  success  of  such  widely-deployed  systems,  as  well  as  provide  a  better  understanding  for  the  research  community  to  build  future  work  on.
■590    ▼aSchool  code:  0033.
■650  4▼aComputer  science
■653    ▼aData  efficiency
■653    ▼aLarge  language  models
■653    ▼aMachine  learning
■653    ▼aRecommender  systems
■690    ▼a0800
■690    ▼a0984
■71020▼aUniversity  of  California,  San  Diego▼bComputer  Science  and  Engineering.
■7730  ▼tDissertations  Abstracts  International▼g86-05B.
■790    ▼a0033
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17162134▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF12887 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.