서브메뉴
검색
Towards Better Language Models: Algorithms, Architectures, and Applications
Towards Better Language Models: Algorithms, Architectures, and Applications
Detailed Information
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211152826
- ISBN
- 9798384090960
- DDC
- 004
- 저자명
- Wu, Qingyang.
- 서명/저자
- Towards Better Language Models: Algorithms, Architectures, and Applications
- 발행사항
- [Sl] : Columbia University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 261 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-03, Section: A.
- 주기사항
- Advisor: Yu, Zhou.
- 학위논문주기
- Thesis (Ph.D.)--Columbia University, 2024.
- 초록/해제
- 요약This thesis explores the advancement of language models by focusing on three important perspectives: Algorithms, Architectures, and Applications. We aim to improve the performance, efficiency, and practical usage of these language models. Specifically, we studied reinforcement learning for language models, recurrent memory-augmented transformers, and practical applications in text generation and dialogue systems.Firstly, we address the limitations of traditional training algorithm maximum likelihood estimation (MLE). We propose TextGAIL, a generative adversarial imitation learning framework, which combines large pre-trained language models with adversarial training to improve the quality and diversity of generated text. We further explore a modern reinforcement learning from human feedback (RLHF) pipeline to more effectively align language model outputs with human preferences.Next, we investigate architecture improvements with Recurrent Memory-Augmented Transformers. In this direction, we first introduce Memformer, an autoregressive model that utilizes an external dynamic memory for efficient long-sequence processing. We build upon Memformer and propose MemBART, a stateful memory-augmented Transformer encoder-decoder model. Recurrent Memory-Augmented Transformers demonstrate superior performance and efficiency in handling long contexts compared to traditional Transformer architectures.Finally, we make several contributions on effectively applying language models to dialogue systems in practice. We design task-oriented dialogue systems that leverage pre-trained language models to significantly reduce the need for human annotations. We also introduce DiactTOD, a novel approach to improving the out of distribution generalization ability of of dialogue act-controlled generation in task-oriented systems. In this thesis we also make progress by expanding the scope of traditional task-oriented dialogue systems by proposing a novel paradigm that utilizes external knowledge tools to provide more accurate knowledge. Our penultimate application tackles the data-scarcity problem common in many real-world dialogue systems. We propose an automatic data augmentation technique to improve training efficacy. Lastly, we make progress on end-user experiences by presenting FaceChat, a multimodal dialogue framework enabling emotionally-sensitive, face-to-face interactions, demonstrating the potential of multimodal language models in various applications.Our work highlights the significance of building better language models, demonstrating how these improvements can positively impact a wide range of downstream tasks and applications. Our work makes a meaningful contribution to language model research, providing valuable insights and methodologies for developing more powerful and efficient models.
- 일반주제명
- Computer science
- 일반주제명
- Language
- 일반주제명
- Computer engineering
- 일반주제명
- Information technology
- 키워드
- Dialog systems
- 키워드
- Language models
- 기타저자
- Columbia University Computer Science
- 기본자료저록
- Dissertations Abstracts International. 86-03A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017164054
■00520250211152826
■006m o d
■007cr#unu||||||||
■020 ▼a9798384090960
■035 ▼a(MiAaPQ)AAI31560134
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aWu, Qingyang.
■24510▼aTowards Better Language Models: Algorithms, Architectures, and Applications
■260 ▼a[Sl]▼bColumbia University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a261 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-03, Section: A.
■500 ▼aAdvisor: Yu, Zhou.
■5021 ▼aThesis (Ph.D.)--Columbia University, 2024.
■520 ▼aThis thesis explores the advancement of language models by focusing on three important perspectives: Algorithms, Architectures, and Applications. We aim to improve the performance, efficiency, and practical usage of these language models. Specifically, we studied reinforcement learning for language models, recurrent memory-augmented transformers, and practical applications in text generation and dialogue systems.Firstly, we address the limitations of traditional training algorithm maximum likelihood estimation (MLE). We propose TextGAIL, a generative adversarial imitation learning framework, which combines large pre-trained language models with adversarial training to improve the quality and diversity of generated text. We further explore a modern reinforcement learning from human feedback (RLHF) pipeline to more effectively align language model outputs with human preferences.Next, we investigate architecture improvements with Recurrent Memory-Augmented Transformers. In this direction, we first introduce Memformer, an autoregressive model that utilizes an external dynamic memory for efficient long-sequence processing. We build upon Memformer and propose MemBART, a stateful memory-augmented Transformer encoder-decoder model. Recurrent Memory-Augmented Transformers demonstrate superior performance and efficiency in handling long contexts compared to traditional Transformer architectures.Finally, we make several contributions on effectively applying language models to dialogue systems in practice. We design task-oriented dialogue systems that leverage pre-trained language models to significantly reduce the need for human annotations. We also introduce DiactTOD, a novel approach to improving the out of distribution generalization ability of of dialogue act-controlled generation in task-oriented systems. In this thesis we also make progress by expanding the scope of traditional task-oriented dialogue systems by proposing a novel paradigm that utilizes external knowledge tools to provide more accurate knowledge. Our penultimate application tackles the data-scarcity problem common in many real-world dialogue systems. We propose an automatic data augmentation technique to improve training efficacy. Lastly, we make progress on end-user experiences by presenting FaceChat, a multimodal dialogue framework enabling emotionally-sensitive, face-to-face interactions, demonstrating the potential of multimodal language models in various applications.Our work highlights the significance of building better language models, demonstrating how these improvements can positively impact a wide range of downstream tasks and applications. Our work makes a meaningful contribution to language model research, providing valuable insights and methodologies for developing more powerful and efficient models.
■590 ▼aSchool code: 0054.
■650 4▼aComputer science
■650 4▼aLanguage
■650 4▼aComputer engineering
■650 4▼aInformation technology
■653 ▼aDialog systems
■653 ▼aLanguage models
■653 ▼aNatural language processing
■653 ▼aReinforcement learning
■653 ▼aRecurrent Memory-Augmented Transformers
■690 ▼a0984
■690 ▼a0489
■690 ▼a0464
■690 ▼a0679
■71020▼aColumbia University▼bComputer Science.
■7730 ▼tDissertations Abstracts International▼g86-03A.
■790 ▼a0054
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164054▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.
Preview
Export
ChatGPT Discussion
AI Recommended Related Books
Подробнее информация.
- Бронирование
- не существует
- моя папка
- Первый запрос зрения
- Non-Book Loan Application
- Nighttime Book Loan Application
Available after logging in.


