본문

서브메뉴

Data Curation for Foundation Model Training
Data Curation for Foundation Model Training  / Georgios Smyrnis
Data Curation for Foundation Model Training

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260311091531.5
ISBN  
9798270231736
DDC  
006.31
저자명  
Smyrnis, Georgios
서명/저자  
Data Curation for Foundation Model Training / Georgios Smyrnis
발행사항  
[Sl] : The University of Texas at Austin, 2025
형태사항  
1 electronic resource (286 pages)
주기사항  
Source: Dissertations Abstracts International, Volume: 87-06, Section: A.
주기사항  
Advisors: Dimakis, Alexandros G. Committee members: Schmidt, Ludwig; Sanghavi, Sujay; Shakkottai, Sanjay; Tamir, Jonathan.
학위논문주기  
- Ph.D. : The University of Texas at Austin, 2025.
초록/해제  
요약Machine learning research has historically focused on algorithmic improvements, with better training methods driving innovation. Given that the amount of data available for training these models was often limited, research aimed on improving the way these relatively small amounts of data could be used. More recently, this focus has shifted from iteration on the algorithms to iteration on the data itself. Techniques such as data filtering, reannotation and data mixing, among many others, are often used to improve upon the data itself.In this work, we will examine ways to increase the quality of datasets in image-text and language domains, with a focus on dataset curation and filtering. Given the amount of readily available data on the web, these techniques can be reliably applied as methods to increase the downstream performance of models, with the improvement stemming directly from the higher quality datasets they were trained on. We will also examine dataset curation from the view of synthetic dataset generation in the domain of language model fine-tuning. These synthetic datasets can be used to distill reasoning capabilities from large, high-performing reasoning models to smaller, more compact ones, improving their usability and reducing inference costs.
언어주기  
English
일반주제명  
Linguistics
일반주제명  
Computer science
일반주제명  
Information technology
키워드  
Machine learning
키워드  
Synthetic datasets
키워드  
Data mixing
키워드  
Downstream performance
기타저자  
The University of Texas at Austin Electrical and Computer Engineering
기본자료저록  
Dissertations Abstracts International. 87-06A.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260311s2025        us                                    eng  d
■001000017361189
■00520260311091531.5
■006m          o    d                
■007cr|nu||||||||
■020    ▼a9798270231736
■040    ▼aMiAaPQD▼beng▼cMiAaPQD▼erda
■082    ▼a006.31
■1001  ▼aSmyrnis,  Georgios▼eauthor.
■24510▼aData  Curation  for  Foundation  Model  Training  ▼cGeorgios  Smyrnis
■260    ▼a[Sl]▼bThe  University  of  Texas  at  Austin▼c2025
■264  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a1  electronic  resource  (286  pages)
■336    ▼atext▼btxt▼2rdacontent
■337    ▼acomputer▼bc▼2rdamedia
■338    ▼aonline  resource▼bcr▼2rdacarrier
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-06,  Section:  A.
■500    ▼aAdvisors:  Dimakis,  Alexandros  G.    Committee  members:  Schmidt,  Ludwig;  Sanghavi,  Sujay;  Shakkottai,  Sanjay;  Tamir,  Jonathan.
■5021  ▼bPh.D.▼cThe  University  of  Texas  at  Austin▼d2025.
■520    ▼aMachine  learning  research  has  historically  focused  on  algorithmic  improvements,  with  better  training  methods  driving  innovation.  Given  that  the  amount  of  data  available  for  training  these  models  was  often  limited,  research  aimed  on  improving  the  way  these  relatively  small  amounts  of  data  could  be  used.  More  recently,  this  focus  has  shifted  from  iteration  on  the  algorithms  to  iteration  on  the  data  itself.  Techniques  such  as  data  filtering,  reannotation  and  data  mixing,  among  many  others,  are  often  used  to  improve  upon  the  data  itself.In  this  work,  we  will  examine  ways  to  increase  the  quality  of  datasets  in  image-text  and  language  domains,  with  a  focus  on  dataset  curation  and  filtering.  Given  the  amount  of  readily  available  data  on  the  web,  these  techniques  can  be  reliably  applied  as  methods  to  increase  the  downstream  performance  of  models,  with  the  improvement  stemming  directly  from  the  higher  quality  datasets  they  were  trained  on.  We  will  also  examine  dataset  curation  from  the  view  of  synthetic  dataset  generation  in  the  domain  of  language  model  fine-tuning.  These  synthetic  datasets  can  be  used  to  distill  reasoning  capabilities  from  large,  high-performing  reasoning  models  to  smaller,  more  compact  ones,  improving  their  usability  and  reducing  inference  costs.
■546    ▼aEnglish
■590    ▼aSchool  code:  0227
■650  4▼aLinguistics
■650  4▼aComputer  science
■650  4▼aInformation  technology
■653    ▼aMachine  learning
■653    ▼aSynthetic  datasets
■653    ▼aData  mixing
■653    ▼aDownstream  performance
■7102  ▼aThe  University  of  Texas  at  Austin▼bElectrical  and  Computer  Engineering.▼edegree  granting  institution.
■7201  ▼aDimakis,  Alexandros  G.▼edegree  supervisor.
■7730  ▼tDissertations  Abstracts  International▼g87-06A.
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17361189▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF17816 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.