서브메뉴
검색
Data Curation for Foundation Model Training
Data Curation for Foundation Model Training
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260311091531.5
- ISBN
- 9798270231736
- DDC
- 006.31
- 서명/저자
- Data Curation for Foundation Model Training / Georgios Smyrnis
- 발행사항
- [Sl] : The University of Texas at Austin, 2025
- 형태사항
- 1 electronic resource (286 pages)
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-06, Section: A.
- 주기사항
- Advisors: Dimakis, Alexandros G. Committee members: Schmidt, Ludwig; Sanghavi, Sujay; Shakkottai, Sanjay; Tamir, Jonathan.
- 학위논문주기
- - Ph.D. : The University of Texas at Austin, 2025.
- 초록/해제
- 요약Machine learning research has historically focused on algorithmic improvements, with better training methods driving innovation. Given that the amount of data available for training these models was often limited, research aimed on improving the way these relatively small amounts of data could be used. More recently, this focus has shifted from iteration on the algorithms to iteration on the data itself. Techniques such as data filtering, reannotation and data mixing, among many others, are often used to improve upon the data itself.In this work, we will examine ways to increase the quality of datasets in image-text and language domains, with a focus on dataset curation and filtering. Given the amount of readily available data on the web, these techniques can be reliably applied as methods to increase the downstream performance of models, with the improvement stemming directly from the higher quality datasets they were trained on. We will also examine dataset curation from the view of synthetic dataset generation in the domain of language model fine-tuning. These synthetic datasets can be used to distill reasoning capabilities from large, high-performing reasoning models to smaller, more compact ones, improving their usability and reducing inference costs.
- 언어주기
- English
- 일반주제명
- Linguistics
- 일반주제명
- Computer science
- 일반주제명
- Information technology
- 키워드
- Machine learning
- 키워드
- Data mixing
- 기타저자
- The University of Texas at Austin Electrical and Computer Engineering
- 기본자료저록
- Dissertations Abstracts International. 87-06A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260311s2025 us eng d■001000017361189
■00520260311091531.5
■006m o d
■007cr|nu||||||||
■020 ▼a9798270231736
■040 ▼aMiAaPQD▼beng▼cMiAaPQD▼erda
■082 ▼a006.31
■1001 ▼aSmyrnis, Georgios▼eauthor.
■24510▼aData Curation for Foundation Model Training ▼cGeorgios Smyrnis
■260 ▼a[Sl]▼bThe University of Texas at Austin▼c2025
■264 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a1 electronic resource (286 pages)
■336 ▼atext▼btxt▼2rdacontent
■337 ▼acomputer▼bc▼2rdamedia
■338 ▼aonline resource▼bcr▼2rdacarrier
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-06, Section: A.
■500 ▼aAdvisors: Dimakis, Alexandros G. Committee members: Schmidt, Ludwig; Sanghavi, Sujay; Shakkottai, Sanjay; Tamir, Jonathan.
■5021 ▼bPh.D.▼cThe University of Texas at Austin▼d2025.
■520 ▼aMachine learning research has historically focused on algorithmic improvements, with better training methods driving innovation. Given that the amount of data available for training these models was often limited, research aimed on improving the way these relatively small amounts of data could be used. More recently, this focus has shifted from iteration on the algorithms to iteration on the data itself. Techniques such as data filtering, reannotation and data mixing, among many others, are often used to improve upon the data itself.In this work, we will examine ways to increase the quality of datasets in image-text and language domains, with a focus on dataset curation and filtering. Given the amount of readily available data on the web, these techniques can be reliably applied as methods to increase the downstream performance of models, with the improvement stemming directly from the higher quality datasets they were trained on. We will also examine dataset curation from the view of synthetic dataset generation in the domain of language model fine-tuning. These synthetic datasets can be used to distill reasoning capabilities from large, high-performing reasoning models to smaller, more compact ones, improving their usability and reducing inference costs.
■546 ▼aEnglish
■590 ▼aSchool code: 0227
■650 4▼aLinguistics
■650 4▼aComputer science
■650 4▼aInformation technology
■653 ▼aMachine learning
■653 ▼aSynthetic datasets
■653 ▼aData mixing
■653 ▼aDownstream performance
■7102 ▼aThe University of Texas at Austin▼bElectrical and Computer Engineering.▼edegree granting institution.
■7201 ▼aDimakis, Alexandros G.▼edegree supervisor.
■7730 ▼tDissertations Abstracts International▼g87-06A.
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17361189▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


