본문

서브메뉴

Fusing Multimodal Knowledge in Language Models
Fusing Multimodal Knowledge in Language Models
Fusing Multimodal Knowledge in Language Models

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211151403
ISBN  
9798382230481
DDC  
004
저자명  
Michihiro Yasunaga.
서명/저자  
Fusing Multimodal Knowledge in Language Models
발행사항  
[Sl] : Stanford University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
226 p
주기사항  
Source: Dissertations Abstracts International, Volume: 85-11, Section: B.
주기사항  
Advisor: Jure Leskovec;Percy Liang.
학위논문주기  
Thesis (Ph.D.)--Stanford University, 2024.
초록/해제  
요약Language models, such as GPT-4, have the capability to generate textual responses to user queries. They are used across various tasks, including question answering, translation, summarization, and personal assistance. However, to create more versatile AI assistants, these models need to handle more diverse and complex tasks involving domain or visual knowledge, such as answering medical questions and explaining or generating images. This necessity motivates the development of models that can access and leverage diverse knowledge sources beyond text, such as databases and images.In this thesis, we aim to develop language models capable of using multimodal knowledge, encompassing text, knowledge graphs, and images, to address various user queries. Text provides broad and contextually rich knowledge, knowledge graphs often supply structured domain knowledge, and images facilitate various visual applications.This thesis consists of five chapters. The first chapter introduces methods for language models to efficiently learn knowledge from textual data. Specifically, we train language models on a sequence of multiple related documents, encouraging them to learn and reason about knowledge with long-range dependencies. This approach yields strong performance on complex long-context and multi-step reasoning tasks. In the second chapter, we introduce methods that enable language models to harness knowledge graph information. Specifically, we develop a new model architecture, a hybrid of language models and graph neural networks, along with a training objective that fuses text and knowledge graph representations. This method demonstrates strong performance on tasks involving domain knowledge, such as medical question answering. In the third chapter, to empower language models to use and generate visual content alongside textual information, we design unified multimodal models capable of encoding, retrieving, and decoding interleaved sequences of text and images. The model employs a retriever to fetch textual or visual knowledge and integrates it into a multimodal Transformer that encodes and decodes both text and images using token representations. Finally, in the forth and fifth chapters, we demonstrate the application of textual, structured, and visual knowledge fusion techniques to solve practical healthcare tasks, including clinical trial outcome prediction and multimodal medical question answering.In summary, this thesis builds models capable of comprehending and generating multimodal content, spanning text, knowledge graphs, and images.
일반주제명  
Computer science
키워드  
Database
키워드  
Language models
키워드  
Summarization
키워드  
AI assistants
기타저자  
Stanford University.
기본자료저록  
Dissertations Abstracts International. 85-11B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017161489
■00520250211151403
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798382230481
■035    ▼a(MiAaPQ)AAI31255788
■035    ▼a(MiAaPQ)dz688yd5162
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aMichihiro  Yasunaga.
■24510▼aFusing  Multimodal  Knowledge  in  Language  Models
■260    ▼a[Sl]▼bStanford  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a226  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  85-11,  Section:  B.
■500    ▼aAdvisor:  Jure  Leskovec;Percy  Liang.
■5021  ▼aThesis  (Ph.D.)--Stanford  University,  2024.
■520    ▼aLanguage  models,  such  as  GPT-4,  have  the  capability  to  generate  textual  responses  to  user  queries.  They  are  used  across  various  tasks,  including  question  answering,  translation,  summarization,  and  personal  assistance.  However,  to  create  more  versatile  AI  assistants,  these  models  need  to  handle  more  diverse  and  complex  tasks  involving  domain  or  visual  knowledge,  such  as  answering  medical  questions  and  explaining  or  generating  images.  This  necessity  motivates  the  development  of  models  that  can  access  and  leverage  diverse  knowledge  sources  beyond  text,  such  as  databases  and  images.In  this  thesis,  we  aim  to  develop  language  models  capable  of  using  multimodal  knowledge,  encompassing  text,  knowledge  graphs,  and  images,  to  address  various  user  queries.  Text  provides  broad  and  contextually  rich  knowledge,  knowledge  graphs  often  supply  structured  domain  knowledge,  and  images  facilitate  various  visual  applications.This  thesis  consists  of  five  chapters.  The  first  chapter  introduces  methods  for  language  models  to  efficiently  learn  knowledge  from  textual  data.  Specifically,  we  train  language  models  on  a  sequence  of  multiple  related  documents,  encouraging  them  to  learn  and  reason  about  knowledge  with  long-range  dependencies.  This  approach  yields  strong  performance  on  complex  long-context  and  multi-step  reasoning  tasks.  In  the  second  chapter,  we  introduce  methods  that  enable  language  models  to  harness  knowledge  graph  information.  Specifically,  we  develop  a  new  model  architecture,  a  hybrid  of  language  models  and  graph  neural  networks,  along  with  a  training  objective  that  fuses  text  and  knowledge  graph  representations.  This  method  demonstrates  strong  performance  on  tasks  involving  domain  knowledge,  such  as  medical  question  answering.  In  the  third  chapter,  to  empower  language  models  to  use  and  generate  visual  content  alongside  textual  information,  we  design  unified  multimodal  models  capable  of  encoding,  retrieving,  and  decoding  interleaved  sequences  of  text  and  images.  The  model  employs  a  retriever  to  fetch  textual  or  visual  knowledge  and  integrates  it  into  a  multimodal  Transformer  that  encodes  and  decodes  both  text  and  images  using  token  representations.  Finally,  in  the  forth  and  fifth  chapters,  we  demonstrate  the  application  of  textual,  structured,  and  visual  knowledge  fusion  techniques  to  solve  practical  healthcare  tasks,  including  clinical  trial  outcome  prediction  and  multimodal  medical  question  answering.In  summary,  this  thesis  builds  models  capable  of  comprehending  and  generating  multimodal  content,  spanning  text,  knowledge  graphs,  and  images.
■590    ▼aSchool  code:  0212.
■650  4▼aComputer  science
■653    ▼aDatabase
■653    ▼aLanguage  models
■653    ▼aSummarization
■653    ▼aAI  assistants
■690    ▼a0984
■690    ▼a0800
■71020▼aStanford  University.
■7730  ▼tDissertations  Abstracts  International▼g85-11B.
■790    ▼a0212
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17161489▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF12967 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.