본문

서브메뉴

Towards Better Language Models: Algorithms, Architectures, and Applications
Towards Better Language Models: Algorithms, Architectures, and Applications
Towards Better Language Models: Algorithms, Architectures, and Applications

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211152826
ISBN  
9798384090960
DDC  
004
저자명  
Wu, Qingyang.
서명/저자  
Towards Better Language Models: Algorithms, Architectures, and Applications
발행사항  
[Sl] : Columbia University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
261 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-03, Section: A.
주기사항  
Advisor: Yu, Zhou.
학위논문주기  
Thesis (Ph.D.)--Columbia University, 2024.
초록/해제  
요약This thesis explores the advancement of language models by focusing on three important perspectives: Algorithms, Architectures, and Applications. We aim to improve the performance, efficiency, and practical usage of these language models. Specifically, we studied reinforcement learning for language models, recurrent memory-augmented transformers, and practical applications in text generation and dialogue systems.Firstly, we address the limitations of traditional training algorithm maximum likelihood estimation (MLE). We propose TextGAIL, a generative adversarial imitation learning framework, which combines large pre-trained language models with adversarial training to improve the quality and diversity of generated text. We further explore a modern reinforcement learning from human feedback (RLHF) pipeline to more effectively align language model outputs with human preferences.Next, we investigate architecture improvements with Recurrent Memory-Augmented Transformers. In this direction, we first introduce Memformer, an autoregressive model that utilizes an external dynamic memory for efficient long-sequence processing. We build upon Memformer and propose MemBART, a stateful memory-augmented Transformer encoder-decoder model. Recurrent Memory-Augmented Transformers demonstrate superior performance and efficiency in handling long contexts compared to traditional Transformer architectures.Finally, we make several contributions on effectively applying language models to dialogue systems in practice. We design task-oriented dialogue systems that leverage pre-trained language models to significantly reduce the need for human annotations. We also introduce DiactTOD, a novel approach to improving the out of distribution generalization ability of of dialogue act-controlled generation in task-oriented systems. In this thesis we also make progress by expanding the scope of traditional task-oriented dialogue systems by proposing a novel paradigm that utilizes external knowledge tools to provide more accurate knowledge. Our penultimate application tackles the data-scarcity problem common in many real-world dialogue systems. We propose an automatic data augmentation technique to improve training efficacy. Lastly, we make progress on end-user experiences by presenting FaceChat, a multimodal dialogue framework enabling emotionally-sensitive, face-to-face interactions, demonstrating the potential of multimodal language models in various applications.Our work highlights the significance of building better language models, demonstrating how these improvements can positively impact a wide range of downstream tasks and applications. Our work makes a meaningful contribution to language model research, providing valuable insights and methodologies for developing more powerful and efficient models.
일반주제명  
Computer science
일반주제명  
Language
일반주제명  
Computer engineering
일반주제명  
Information technology
키워드  
Dialog systems
키워드  
Language models
키워드  
Natural language processing
키워드  
Reinforcement learning
키워드  
Recurrent Memory-Augmented Transformers
기타저자  
Columbia University Computer Science
기본자료저록  
Dissertations Abstracts International. 86-03A.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017164054
■00520250211152826
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798384090960
■035    ▼a(MiAaPQ)AAI31560134
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aWu,  Qingyang.
■24510▼aTowards  Better  Language  Models:  Algorithms,  Architectures,  and  Applications
■260    ▼a[Sl]▼bColumbia  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a261  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-03,  Section:  A.
■500    ▼aAdvisor:  Yu,  Zhou.
■5021  ▼aThesis  (Ph.D.)--Columbia  University,  2024.
■520    ▼aThis  thesis  explores  the  advancement  of  language  models  by  focusing  on  three  important  perspectives:  Algorithms,  Architectures,  and  Applications.  We  aim  to  improve  the  performance,  efficiency,  and  practical  usage  of  these  language  models.  Specifically,  we  studied  reinforcement  learning  for  language  models,  recurrent  memory-augmented  transformers,  and  practical  applications  in  text  generation  and  dialogue  systems.Firstly,  we  address  the  limitations  of  traditional  training  algorithm  maximum  likelihood  estimation  (MLE).  We  propose  TextGAIL,  a  generative  adversarial  imitation  learning  framework,  which  combines  large  pre-trained  language  models  with  adversarial  training  to  improve  the  quality  and  diversity  of  generated  text.  We  further  explore  a  modern  reinforcement  learning  from  human  feedback  (RLHF)  pipeline  to  more  effectively  align  language  model  outputs  with  human  preferences.Next,  we  investigate  architecture  improvements  with  Recurrent  Memory-Augmented  Transformers.  In  this  direction,  we  first  introduce  Memformer,  an  autoregressive  model  that  utilizes  an  external  dynamic  memory  for  efficient  long-sequence  processing.  We  build  upon  Memformer  and  propose  MemBART,  a  stateful  memory-augmented  Transformer  encoder-decoder  model.  Recurrent  Memory-Augmented  Transformers  demonstrate  superior  performance  and  efficiency  in  handling  long  contexts  compared  to  traditional  Transformer  architectures.Finally,  we  make  several  contributions  on  effectively  applying  language  models  to  dialogue  systems  in  practice.  We  design  task-oriented  dialogue  systems  that  leverage  pre-trained  language  models  to  significantly  reduce  the  need  for  human  annotations.  We  also  introduce  DiactTOD,  a  novel  approach  to  improving  the  out  of  distribution  generalization  ability  of  of  dialogue  act-controlled  generation  in  task-oriented  systems.  In  this  thesis  we  also  make  progress  by  expanding  the  scope  of  traditional  task-oriented  dialogue  systems  by  proposing  a  novel  paradigm  that  utilizes  external  knowledge  tools  to  provide  more  accurate  knowledge.  Our  penultimate  application  tackles  the  data-scarcity  problem  common  in  many  real-world  dialogue  systems.  We  propose  an  automatic  data  augmentation  technique  to  improve  training  efficacy.  Lastly,  we  make  progress  on  end-user  experiences  by  presenting  FaceChat,  a  multimodal  dialogue  framework  enabling  emotionally-sensitive,  face-to-face  interactions,  demonstrating  the  potential  of  multimodal  language  models  in  various  applications.Our  work  highlights  the  significance  of  building  better  language  models,  demonstrating  how  these  improvements  can  positively  impact  a  wide  range  of  downstream  tasks  and  applications.  Our  work  makes  a  meaningful  contribution  to  language  model  research,  providing  valuable  insights  and  methodologies  for  developing  more  powerful  and  efficient  models.
■590    ▼aSchool  code:  0054.
■650  4▼aComputer  science
■650  4▼aLanguage
■650  4▼aComputer  engineering
■650  4▼aInformation  technology
■653    ▼aDialog  systems
■653    ▼aLanguage  models
■653    ▼aNatural  language  processing
■653    ▼aReinforcement  learning
■653    ▼aRecurrent  Memory-Augmented  Transformers
■690    ▼a0984
■690    ▼a0489
■690    ▼a0464
■690    ▼a0679
■71020▼aColumbia  University▼bComputer  Science.
■7730  ▼tDissertations  Abstracts  International▼g86-03A.
■790    ▼a0054
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164054▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF14122 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.