서브메뉴
검색
Interpretation Errors: Extracting Functionality From Generative Models of Language by Understanding Them Better- [electronic resource]
Interpretation Errors: Extracting Functionality From Generative Models of Language by Understanding Them Better- [electronic resource]
상세정보
- 자료유형
- 학위논문파일 국외
- 최종처리일시
- 20240214101651
- ISBN
- 9798380328883
- DDC
- 004
- 저자명
- Holtzman, Ari.
- 서명/저자
- Interpretation Errors: Extracting Functionality From Generative Models of Language by Understanding Them Better - [electronic resource]
- 발행사항
- [S.l.]: : University of Washington., 2023
- 발행사항
- Ann Arbor : : ProQuest Dissertations & Theses,, 2023
- 형태사항
- 1 online resource(129 p.)
- 주기사항
- Source: Dissertations Abstracts International, Volume: 85-03, Section: A.
- 주기사항
- Advisor: Zettlemoyer, Luke.
- 학위논문주기
- Thesis (Ph.D.)--University of Washington, 2023.
- 사용제한주기
- This item must not be sold to any third party vendors.
- 초록/해제
- 요약The rise of large language models as the workhorse of NLP, and the continuous release of better models (OpenAI, 2023; Pichai, 2023; Schulman et al., 2022, inter alia) has created a strange situation: we have models that are more powerful language generators than ever before, but since we did not design them for a specific purpose we struggle to understand how they should be used or what their idiosyncracies are.This dissertation describes three empirical projects that sought to characterize the underlying behavior of language models and, importantly, to make them more reliable tools for generating and selecting text where this behavior does not match up with the tasks we would like models to complete. Each project attempts to understand what language models and accompanying inference methods currently optimize for, to characterize the gap between that and the true objective of a potential user, and to close it with some new inference method. An emergent theme through these works is that models are already doing what we trained them to do quite well-and it is often the experimenters and practitioners who misunderstand precisely what we trained models to do in the first place. We conclude with a conceptual analysis of how we should study generative models going forward-as models keep improving and new, unanticipated uses and misuses become ever more available.The first half of this dissertation concerns two works, Neural Text Degeneration and Surface Form Competition-two failure modes of generative models that occur when probability is viewed as equivalent to "correctness" in text generation and multiple choice scenarios, respectively. For these works we describe the resultant issues, and propose inference methods that largely alleviate them.The second half of this dissertation goes deeper into the question of how generative models of language capture the communicative goals that humans are optimizing: first with Learning to Write, operationalizing communicative goals into auxiliary search objectives for text decoding, and then with Generative Models as a Complex Systems Science, which presents a framework to think about the study of generative models as NLP shifts to analyzing systems that are often infeasible to replicate.How does a model that is predicting the distribution of next tokens understand-and fail to understand-the structure of an essay? This is precisely the kind of question we must face head-on in the new science of generative models.
- 일반주제명
- Computer science.
- 일반주제명
- Computer engineering.
- 일반주제명
- Linguistics.
- 키워드
- Language models
- 키워드
- Text decoding
- 기타저자
- University of Washington Computer Science and Engineering
- 기본자료저록
- Dissertations Abstracts International. 85-03A.
- 기본자료저록
- Dissertation Abstract International
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008240612s2023 us |||||||||||||||c||eng d■001000016934766
■00520240214101651
■006m o d
■007cr#unu||||||||
■020 ▼a9798380328883
■035 ▼a(MiAaPQ)AAI30634344
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aHoltzman, Ari.
■24510▼aInterpretation Errors: Extracting Functionality From Generative Models of Language by Understanding Them Better▼h[electronic resource]
■260 ▼a[S.l.]:▼bUniversity of Washington. ▼c2023
■260 1▼aAnn Arbor :▼bProQuest Dissertations & Theses, ▼c2023
■300 ▼a1 online resource(129 p.)
■500 ▼aSource: Dissertations Abstracts International, Volume: 85-03, Section: A.
■500 ▼aAdvisor: Zettlemoyer, Luke.
■5021 ▼aThesis (Ph.D.)--University of Washington, 2023.
■506 ▼aThis item must not be sold to any third party vendors.
■520 ▼aThe rise of large language models as the workhorse of NLP, and the continuous release of better models (OpenAI, 2023; Pichai, 2023; Schulman et al., 2022, inter alia) has created a strange situation: we have models that are more powerful language generators than ever before, but since we did not design them for a specific purpose we struggle to understand how they should be used or what their idiosyncracies are.This dissertation describes three empirical projects that sought to characterize the underlying behavior of language models and, importantly, to make them more reliable tools for generating and selecting text where this behavior does not match up with the tasks we would like models to complete. Each project attempts to understand what language models and accompanying inference methods currently optimize for, to characterize the gap between that and the true objective of a potential user, and to close it with some new inference method. An emergent theme through these works is that models are already doing what we trained them to do quite well-and it is often the experimenters and practitioners who misunderstand precisely what we trained models to do in the first place. We conclude with a conceptual analysis of how we should study generative models going forward-as models keep improving and new, unanticipated uses and misuses become ever more available.The first half of this dissertation concerns two works, Neural Text Degeneration and Surface Form Competition-two failure modes of generative models that occur when probability is viewed as equivalent to "correctness" in text generation and multiple choice scenarios, respectively. For these works we describe the resultant issues, and propose inference methods that largely alleviate them.The second half of this dissertation goes deeper into the question of how generative models of language capture the communicative goals that humans are optimizing: first with Learning to Write, operationalizing communicative goals into auxiliary search objectives for text decoding, and then with Generative Models as a Complex Systems Science, which presents a framework to think about the study of generative models as NLP shifts to analyzing systems that are often infeasible to replicate.How does a model that is predicting the distribution of next tokens understand-and fail to understand-the structure of an essay? This is precisely the kind of question we must face head-on in the new science of generative models.
■590 ▼aSchool code: 0250.
■650 4▼aComputer science.
■650 4▼aComputer engineering.
■650 4▼aLinguistics.
■653 ▼aLanguage models
■653 ▼aCommunicative goals
■653 ▼aText decoding
■653 ▼aAnalyzing systems
■653 ▼aInterpretation errors
■690 ▼a0984
■690 ▼a0464
■690 ▼a0290
■71020▼aUniversity of Washington▼bComputer Science and Engineering.
■7730 ▼tDissertations Abstracts International▼g85-03A.
■773 ▼tDissertation Abstract International
■790 ▼a0250
■791 ▼aPh.D.
■792 ▼a2023
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T16934766▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.
■980 ▼a202402▼f2024


