서브메뉴
검색
Multimodal Reasoning With Fine-Grained Knowledge Representation
Multimodal Reasoning With Fine-Grained Knowledge Representation
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211152934
- ISBN
- 9798384486329
- DDC
- 004
- 저자명
- Wang, Zhecan.
- 서명/저자
- Multimodal Reasoning With Fine-Grained Knowledge Representation
- 발행사항
- [Sl] : Columbia University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 135 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-04, Section: B.
- 주기사항
- Advisor: Chang, Shih-Fu.
- 학위논문주기
- Thesis (Ph.D.)--Columbia University, 2024.
- 초록/해제
- 요약Multimodal reasoning, especially the processing that involves commonsense, is a vital capability for humans, encompassing a wide range of practical applications, from understanding visual cues while driving to interpreting emotions and intentions in social interactions and efficiently planning and executing household chores. Therefore, multimodal (commonsense) reasoning represents an important step when developing advanced AI systems that aim to imitate human-level capabilities. However, existing methods struggle to achieve this due to several factors, including the under-utilization of fine-grained multimodal information, lack of transparency, and unexplainable and unreliable behaviors of the AI models.Our research addresses these challenges by focusing on improving AI models' multimodal (commonsense) reasoning through the utilization of fine-grained knowledge representation. We begin by developing transformer-based models to extract fine-grained knowledge across various modalities. We then propose novel solutions to leverage this extracted knowledge to enhance AI models' learning of multimodal reasoning, particularly in downstream vision-language understanding tasks such as visual question answering and visual entailment.Beyond the focus on high performance, we further propose approaches that exploit fine-grained multimodal knowledge to enhance our understanding of AI models, thereby improving the explainability of how vision-language models work during multimodal (commonsense) reasoning. Finally, we develop new methods to utilize fine-grained knowledge to create generalized and challenging multimodal benchmarks, designed specifically to evaluate future AI models on their multimodal reasoning capabilities.Throughout our research, we conduct extensive experiments to demonstrate the effectiveness of utilizing fine-grained knowledge in improving AI models for multimodal reasoning. Our work focuses on four key perspectives: knowledge extraction, model learning, explainability, and evaluation, providing a comprehensive approach to advancing the field of AI multimodal reasoning.
- 일반주제명
- Computer science
- 일반주제명
- Computer engineering
- 키워드
- Reasoning
- 키워드
- Machine learning
- 기타저자
- Columbia University Computer Science
- 기본자료저록
- Dissertations Abstracts International. 86-04B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017164215
■00520250211152934
■006m o d
■007cr#unu||||||||
■020 ▼a9798384486329
■035 ▼a(MiAaPQ)AAI31563189
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aWang, Zhecan.
■24510▼aMultimodal Reasoning With Fine-Grained Knowledge Representation
■260 ▼a[Sl]▼bColumbia University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a135 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-04, Section: B.
■500 ▼aAdvisor: Chang, Shih-Fu.
■5021 ▼aThesis (Ph.D.)--Columbia University, 2024.
■520 ▼aMultimodal reasoning, especially the processing that involves commonsense, is a vital capability for humans, encompassing a wide range of practical applications, from understanding visual cues while driving to interpreting emotions and intentions in social interactions and efficiently planning and executing household chores. Therefore, multimodal (commonsense) reasoning represents an important step when developing advanced AI systems that aim to imitate human-level capabilities. However, existing methods struggle to achieve this due to several factors, including the under-utilization of fine-grained multimodal information, lack of transparency, and unexplainable and unreliable behaviors of the AI models.Our research addresses these challenges by focusing on improving AI models' multimodal (commonsense) reasoning through the utilization of fine-grained knowledge representation. We begin by developing transformer-based models to extract fine-grained knowledge across various modalities. We then propose novel solutions to leverage this extracted knowledge to enhance AI models' learning of multimodal reasoning, particularly in downstream vision-language understanding tasks such as visual question answering and visual entailment.Beyond the focus on high performance, we further propose approaches that exploit fine-grained multimodal knowledge to enhance our understanding of AI models, thereby improving the explainability of how vision-language models work during multimodal (commonsense) reasoning. Finally, we develop new methods to utilize fine-grained knowledge to create generalized and challenging multimodal benchmarks, designed specifically to evaluate future AI models on their multimodal reasoning capabilities.Throughout our research, we conduct extensive experiments to demonstrate the effectiveness of utilizing fine-grained knowledge in improving AI models for multimodal reasoning. Our work focuses on four key perspectives: knowledge extraction, model learning, explainability, and evaluation, providing a comprehensive approach to advancing the field of AI multimodal reasoning.
■590 ▼aSchool code: 0054.
■650 4▼aComputer science
■650 4▼aComputer engineering
■653 ▼aReasoning
■653 ▼aMachine learning
■653 ▼aMultimodal learning
■653 ▼aMultimodal representation
■653 ▼aVisual entailment
■690 ▼a0800
■690 ▼a0984
■690 ▼a0464
■71020▼aColumbia University▼bComputer Science.
■7730 ▼tDissertations Abstracts International▼g86-04B.
■790 ▼a0054
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164215▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


