본문

서브메뉴

Characterizing the Difficulty of Natural Language Datasets for Machine Learning
Characterizing the Difficulty of Natural Language Datasets for Machine Learning
Characterizing the Difficulty of Natural Language Datasets for Machine Learning

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211152716
ISBN  
9798384053477
DDC  
004
저자명  
Yauney, Gregory.
서명/저자  
Characterizing the Difficulty of Natural Language Datasets for Machine Learning
발행사항  
[Sl] : Cornell University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
179 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-03, Section: B.
주기사항  
Advisor: Mimno, David.
학위논문주기  
Thesis (Ph.D.)--Cornell University, 2024.
초록/해제  
요약Machine learning models can now achieve high performance on many natural language classification tasks. But we currently don't know how well a contemporary large language model will perform on a new task without directly trying it out. What makes a task difficult for machine learning models? We focus on the role of data-both a model's training data and that of downstream tasks-to go beyond evaluation performance in characterizing the difficulty of natural language tasks. We intervene throughout the language modeling pipeline, examining the interaction between a task's dataset and a) pretrained representations, b) evaluation, and c) pretraining data. We use random labelings to contextualize the degree of alignment between a task's data and a task's labels under different text representations. We use classifiers that guess uniformly at random, independently across examples, to contextualize a language model's performance on the small datasets typically used to evaluate in-context learning capabilities. We also examine the extent of evidence for the hypothesis that a downstream dataset's similarity to a model's pretraining dataset determines the model's performance. Finally, we turn to case studies across image-text grounding, literary history, and architectural history where we are specifically interested in a model's performance on a given challenging dataset. Understanding the interaction between data and model will make our models ever more reliable on datasets that we care about, ultimately meeting text datasets where they are.
일반주제명  
Computer science
일반주제명  
Computer engineering
키워드  
Classification tasks
키워드  
Machine learning
키워드  
Natural language
키워드  
Text dataset
기타저자  
Cornell University Computer Science
기본자료저록  
Dissertations Abstracts International. 86-03B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017163499
■00520250211152716
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798384053477
■035    ▼a(MiAaPQ)AAI31489088
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aYauney,  Gregory.▼0(orcid)0009-0001-9087-0901
■24510▼aCharacterizing  the  Difficulty  of  Natural  Language  Datasets  for  Machine  Learning
■260    ▼a[Sl]▼bCornell  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a179  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-03,  Section:  B.
■500    ▼aAdvisor:  Mimno,  David.
■5021  ▼aThesis  (Ph.D.)--Cornell  University,  2024.
■520    ▼aMachine  learning  models  can  now  achieve  high  performance  on  many  natural  language  classification  tasks.  But  we  currently  don't  know  how  well  a  contemporary  large  language  model  will  perform  on  a  new  task  without  directly  trying  it  out.  What  makes  a  task  difficult  for  machine  learning  models?  We  focus  on  the  role  of  data-both  a  model's  training  data  and  that  of  downstream  tasks-to  go  beyond  evaluation  performance  in  characterizing  the  difficulty  of  natural  language  tasks.  We  intervene  throughout  the  language  modeling  pipeline,  examining  the  interaction  between  a  task's  dataset  and  a)  pretrained  representations,  b)  evaluation,  and  c)  pretraining  data.  We  use  random  labelings  to  contextualize  the  degree  of  alignment  between  a  task's  data  and  a  task's  labels  under  different  text  representations.  We  use  classifiers  that  guess  uniformly  at  random,  independently  across  examples,  to  contextualize  a  language  model's  performance  on  the  small  datasets  typically  used  to  evaluate  in-context  learning  capabilities.  We  also  examine  the  extent  of  evidence  for  the  hypothesis  that  a  downstream  dataset's  similarity  to  a  model's  pretraining  dataset  determines  the  model's  performance.  Finally,  we  turn  to  case  studies  across  image-text  grounding,  literary  history,  and  architectural  history  where  we  are  specifically  interested  in  a  model's  performance  on  a  given  challenging  dataset.  Understanding  the  interaction  between  data  and  model  will  make  our  models  ever  more  reliable  on  datasets  that  we  care  about,  ultimately  meeting  text  datasets  where  they  are.
■590    ▼aSchool  code:  0058.
■650  4▼aComputer  science
■650  4▼aComputer  engineering
■653    ▼aClassification  tasks
■653    ▼aMachine  learning
■653    ▼aNatural  language
■653    ▼aText  dataset
■690    ▼a0984
■690    ▼a0800
■690    ▼a0464
■71020▼aCornell  University▼bComputer  Science.
■7730  ▼tDissertations  Abstracts  International▼g86-03B.
■790    ▼a0058
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17163499▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF12635 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.