서브메뉴
검색
Characterizing the Difficulty of Natural Language Datasets for Machine Learning
Characterizing the Difficulty of Natural Language Datasets for Machine Learning
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211152716
- ISBN
- 9798384053477
- DDC
- 004
- 저자명
- Yauney, Gregory.
- 서명/저자
- Characterizing the Difficulty of Natural Language Datasets for Machine Learning
- 발행사항
- [Sl] : Cornell University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 179 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-03, Section: B.
- 주기사항
- Advisor: Mimno, David.
- 학위논문주기
- Thesis (Ph.D.)--Cornell University, 2024.
- 초록/해제
- 요약Machine learning models can now achieve high performance on many natural language classification tasks. But we currently don't know how well a contemporary large language model will perform on a new task without directly trying it out. What makes a task difficult for machine learning models? We focus on the role of data-both a model's training data and that of downstream tasks-to go beyond evaluation performance in characterizing the difficulty of natural language tasks. We intervene throughout the language modeling pipeline, examining the interaction between a task's dataset and a) pretrained representations, b) evaluation, and c) pretraining data. We use random labelings to contextualize the degree of alignment between a task's data and a task's labels under different text representations. We use classifiers that guess uniformly at random, independently across examples, to contextualize a language model's performance on the small datasets typically used to evaluate in-context learning capabilities. We also examine the extent of evidence for the hypothesis that a downstream dataset's similarity to a model's pretraining dataset determines the model's performance. Finally, we turn to case studies across image-text grounding, literary history, and architectural history where we are specifically interested in a model's performance on a given challenging dataset. Understanding the interaction between data and model will make our models ever more reliable on datasets that we care about, ultimately meeting text datasets where they are.
- 일반주제명
- Computer science
- 일반주제명
- Computer engineering
- 키워드
- Machine learning
- 키워드
- Natural language
- 키워드
- Text dataset
- 기타저자
- Cornell University Computer Science
- 기본자료저록
- Dissertations Abstracts International. 86-03B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017163499
■00520250211152716
■006m o d
■007cr#unu||||||||
■020 ▼a9798384053477
■035 ▼a(MiAaPQ)AAI31489088
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aYauney, Gregory.▼0(orcid)0009-0001-9087-0901
■24510▼aCharacterizing the Difficulty of Natural Language Datasets for Machine Learning
■260 ▼a[Sl]▼bCornell University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a179 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-03, Section: B.
■500 ▼aAdvisor: Mimno, David.
■5021 ▼aThesis (Ph.D.)--Cornell University, 2024.
■520 ▼aMachine learning models can now achieve high performance on many natural language classification tasks. But we currently don't know how well a contemporary large language model will perform on a new task without directly trying it out. What makes a task difficult for machine learning models? We focus on the role of data-both a model's training data and that of downstream tasks-to go beyond evaluation performance in characterizing the difficulty of natural language tasks. We intervene throughout the language modeling pipeline, examining the interaction between a task's dataset and a) pretrained representations, b) evaluation, and c) pretraining data. We use random labelings to contextualize the degree of alignment between a task's data and a task's labels under different text representations. We use classifiers that guess uniformly at random, independently across examples, to contextualize a language model's performance on the small datasets typically used to evaluate in-context learning capabilities. We also examine the extent of evidence for the hypothesis that a downstream dataset's similarity to a model's pretraining dataset determines the model's performance. Finally, we turn to case studies across image-text grounding, literary history, and architectural history where we are specifically interested in a model's performance on a given challenging dataset. Understanding the interaction between data and model will make our models ever more reliable on datasets that we care about, ultimately meeting text datasets where they are.
■590 ▼aSchool code: 0058.
■650 4▼aComputer science
■650 4▼aComputer engineering
■653 ▼aClassification tasks
■653 ▼aMachine learning
■653 ▼aNatural language
■653 ▼aText dataset
■690 ▼a0984
■690 ▼a0800
■690 ▼a0464
■71020▼aCornell University▼bComputer Science.
■7730 ▼tDissertations Abstracts International▼g86-03B.
■790 ▼a0058
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17163499▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


