서브메뉴
검색
Towards Data-Efficient Machine Learning Systems
Towards Data-Efficient Machine Learning Systems
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211151923
- ISBN
- 9798346738145
- DDC
- 004
- 서명/저자
- Towards Data-Efficient Machine Learning Systems
- 발행사항
- [Sl] : University of California, San Diego, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 172 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-05, Section: B.
- 주기사항
- Advisor: McAuley, Julian.
- 학위논문주기
- Thesis (Ph.D.)--University of California, San Diego, 2024.
- 초록/해제
- 요약The amount of data available to train modern machine learning systems has been increasing rapidly, so much so that we're using, e.g., entirety of the publicly available text data to train state-of-the-art (SoTA) large language models (LLMs), interaction data from billions of users to train SoTA recommender systems, etc. Training of such large machine learning systems on such large datasets entails a high (i) computational runtime, (ii) economical cost, and (iii) carbon footprint; all of which we aim to minimize for different reasons.While a large body of literature develops "model-centric" techniques to better model a given dataset, in this thesis, we develop a "data-centric" viewpoint, where we are interested in techniques that can appropriately summarize a given training dataset, such that models can be trained equally effectively on the data summary vs. training on the much larger original dataset. In addition to being more efficient overall, data-efficient techniques further aim to improve the trained model's quality by stripping away the low-quality and noisy sources of information in the original dataset.More specifically, we develop techniques from two disparate data summarization ideologies: (i) data pruning (a.k.a. coreset construction) techniques that sample the most relevant portions from the dataset using various grounded heuristics, and (ii) data distillation techniques that generate synthetic data-points which summarize the underlying information in the dataset, and are optimized end-to-end using meta-learning. We restrict our scope to training (i) language models on textual datasets, and (ii) recommender systems on user-item interaction datasets.By pushing the frontier of data-efficient training of machine learning systems, we believe our research can effectively contribute to the practical success of such widely-deployed systems, as well as provide a better understanding for the research community to build future work on.
- 일반주제명
- Computer science
- 키워드
- Data efficiency
- 키워드
- Machine learning
- 기타저자
- University of California, San Diego Computer Science and Engineering
- 기본자료저록
- Dissertations Abstracts International. 86-05B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017162134
■00520250211151923
■006m o d
■007cr#unu||||||||
■020 ▼a9798346738145
■035 ▼a(MiAaPQ)AAI31295211
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aSachdeva, Noveen.
■24510▼aTowards Data-Efficient Machine Learning Systems
■260 ▼a[Sl]▼bUniversity of California, San Diego▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a172 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-05, Section: B.
■500 ▼aAdvisor: McAuley, Julian.
■5021 ▼aThesis (Ph.D.)--University of California, San Diego, 2024.
■520 ▼aThe amount of data available to train modern machine learning systems has been increasing rapidly, so much so that we're using, e.g., entirety of the publicly available text data to train state-of-the-art (SoTA) large language models (LLMs), interaction data from billions of users to train SoTA recommender systems, etc. Training of such large machine learning systems on such large datasets entails a high (i) computational runtime, (ii) economical cost, and (iii) carbon footprint; all of which we aim to minimize for different reasons.While a large body of literature develops "model-centric" techniques to better model a given dataset, in this thesis, we develop a "data-centric" viewpoint, where we are interested in techniques that can appropriately summarize a given training dataset, such that models can be trained equally effectively on the data summary vs. training on the much larger original dataset. In addition to being more efficient overall, data-efficient techniques further aim to improve the trained model's quality by stripping away the low-quality and noisy sources of information in the original dataset.More specifically, we develop techniques from two disparate data summarization ideologies: (i) data pruning (a.k.a. coreset construction) techniques that sample the most relevant portions from the dataset using various grounded heuristics, and (ii) data distillation techniques that generate synthetic data-points which summarize the underlying information in the dataset, and are optimized end-to-end using meta-learning. We restrict our scope to training (i) language models on textual datasets, and (ii) recommender systems on user-item interaction datasets.By pushing the frontier of data-efficient training of machine learning systems, we believe our research can effectively contribute to the practical success of such widely-deployed systems, as well as provide a better understanding for the research community to build future work on.
■590 ▼aSchool code: 0033.
■650 4▼aComputer science
■653 ▼aData efficiency
■653 ▼aLarge language models
■653 ▼aMachine learning
■653 ▼aRecommender systems
■690 ▼a0800
■690 ▼a0984
■71020▼aUniversity of California, San Diego▼bComputer Science and Engineering.
■7730 ▼tDissertations Abstracts International▼g86-05B.
■790 ▼a0033
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17162134▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


