서브메뉴
검색
Fusing Multimodal Knowledge in Language Models
Fusing Multimodal Knowledge in Language Models
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211151403
- ISBN
- 9798382230481
- DDC
- 004
- 서명/저자
- Fusing Multimodal Knowledge in Language Models
- 발행사항
- [Sl] : Stanford University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 226 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 85-11, Section: B.
- 주기사항
- Advisor: Jure Leskovec;Percy Liang.
- 학위논문주기
- Thesis (Ph.D.)--Stanford University, 2024.
- 초록/해제
- 요약Language models, such as GPT-4, have the capability to generate textual responses to user queries. They are used across various tasks, including question answering, translation, summarization, and personal assistance. However, to create more versatile AI assistants, these models need to handle more diverse and complex tasks involving domain or visual knowledge, such as answering medical questions and explaining or generating images. This necessity motivates the development of models that can access and leverage diverse knowledge sources beyond text, such as databases and images.In this thesis, we aim to develop language models capable of using multimodal knowledge, encompassing text, knowledge graphs, and images, to address various user queries. Text provides broad and contextually rich knowledge, knowledge graphs often supply structured domain knowledge, and images facilitate various visual applications.This thesis consists of five chapters. The first chapter introduces methods for language models to efficiently learn knowledge from textual data. Specifically, we train language models on a sequence of multiple related documents, encouraging them to learn and reason about knowledge with long-range dependencies. This approach yields strong performance on complex long-context and multi-step reasoning tasks. In the second chapter, we introduce methods that enable language models to harness knowledge graph information. Specifically, we develop a new model architecture, a hybrid of language models and graph neural networks, along with a training objective that fuses text and knowledge graph representations. This method demonstrates strong performance on tasks involving domain knowledge, such as medical question answering. In the third chapter, to empower language models to use and generate visual content alongside textual information, we design unified multimodal models capable of encoding, retrieving, and decoding interleaved sequences of text and images. The model employs a retriever to fetch textual or visual knowledge and integrates it into a multimodal Transformer that encodes and decodes both text and images using token representations. Finally, in the forth and fifth chapters, we demonstrate the application of textual, structured, and visual knowledge fusion techniques to solve practical healthcare tasks, including clinical trial outcome prediction and multimodal medical question answering.In summary, this thesis builds models capable of comprehending and generating multimodal content, spanning text, knowledge graphs, and images.
- 일반주제명
- Computer science
- 키워드
- Database
- 키워드
- Language models
- 키워드
- Summarization
- 키워드
- AI assistants
- 기타저자
- Stanford University.
- 기본자료저록
- Dissertations Abstracts International. 85-11B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017161489
■00520250211151403
■006m o d
■007cr#unu||||||||
■020 ▼a9798382230481
■035 ▼a(MiAaPQ)AAI31255788
■035 ▼a(MiAaPQ)dz688yd5162
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aMichihiro Yasunaga.
■24510▼aFusing Multimodal Knowledge in Language Models
■260 ▼a[Sl]▼bStanford University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a226 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 85-11, Section: B.
■500 ▼aAdvisor: Jure Leskovec;Percy Liang.
■5021 ▼aThesis (Ph.D.)--Stanford University, 2024.
■520 ▼aLanguage models, such as GPT-4, have the capability to generate textual responses to user queries. They are used across various tasks, including question answering, translation, summarization, and personal assistance. However, to create more versatile AI assistants, these models need to handle more diverse and complex tasks involving domain or visual knowledge, such as answering medical questions and explaining or generating images. This necessity motivates the development of models that can access and leverage diverse knowledge sources beyond text, such as databases and images.In this thesis, we aim to develop language models capable of using multimodal knowledge, encompassing text, knowledge graphs, and images, to address various user queries. Text provides broad and contextually rich knowledge, knowledge graphs often supply structured domain knowledge, and images facilitate various visual applications.This thesis consists of five chapters. The first chapter introduces methods for language models to efficiently learn knowledge from textual data. Specifically, we train language models on a sequence of multiple related documents, encouraging them to learn and reason about knowledge with long-range dependencies. This approach yields strong performance on complex long-context and multi-step reasoning tasks. In the second chapter, we introduce methods that enable language models to harness knowledge graph information. Specifically, we develop a new model architecture, a hybrid of language models and graph neural networks, along with a training objective that fuses text and knowledge graph representations. This method demonstrates strong performance on tasks involving domain knowledge, such as medical question answering. In the third chapter, to empower language models to use and generate visual content alongside textual information, we design unified multimodal models capable of encoding, retrieving, and decoding interleaved sequences of text and images. The model employs a retriever to fetch textual or visual knowledge and integrates it into a multimodal Transformer that encodes and decodes both text and images using token representations. Finally, in the forth and fifth chapters, we demonstrate the application of textual, structured, and visual knowledge fusion techniques to solve practical healthcare tasks, including clinical trial outcome prediction and multimodal medical question answering.In summary, this thesis builds models capable of comprehending and generating multimodal content, spanning text, knowledge graphs, and images.
■590 ▼aSchool code: 0212.
■650 4▼aComputer science
■653 ▼aDatabase
■653 ▼aLanguage models
■653 ▼aSummarization
■653 ▼aAI assistants
■690 ▼a0984
■690 ▼a0800
■71020▼aStanford University.
■7730 ▼tDissertations Abstracts International▼g85-11B.
■790 ▼a0212
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17161489▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


