본문

서브메뉴

Multimodal Reasoning With Fine-Grained Knowledge Representation
Multimodal Reasoning With Fine-Grained Knowledge Representation
Multimodal Reasoning With Fine-Grained Knowledge Representation

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211152934
ISBN  
9798384486329
DDC  
004
저자명  
Wang, Zhecan.
서명/저자  
Multimodal Reasoning With Fine-Grained Knowledge Representation
발행사항  
[Sl] : Columbia University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
135 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-04, Section: B.
주기사항  
Advisor: Chang, Shih-Fu.
학위논문주기  
Thesis (Ph.D.)--Columbia University, 2024.
초록/해제  
요약Multimodal reasoning, especially the processing that involves commonsense, is a vital capability for humans, encompassing a wide range of practical applications, from understanding visual cues while driving to interpreting emotions and intentions in social interactions and efficiently planning and executing household chores. Therefore, multimodal (commonsense) reasoning represents an important step when developing advanced AI systems that aim to imitate human-level capabilities. However, existing methods struggle to achieve this due to several factors, including the under-utilization of fine-grained multimodal information, lack of transparency, and unexplainable and unreliable behaviors of the AI models.Our research addresses these challenges by focusing on improving AI models' multimodal (commonsense) reasoning through the utilization of fine-grained knowledge representation. We begin by developing transformer-based models to extract fine-grained knowledge across various modalities. We then propose novel solutions to leverage this extracted knowledge to enhance AI models' learning of multimodal reasoning, particularly in downstream vision-language understanding tasks such as visual question answering and visual entailment.Beyond the focus on high performance, we further propose approaches that exploit fine-grained multimodal knowledge to enhance our understanding of AI models, thereby improving the explainability of how vision-language models work during multimodal (commonsense) reasoning. Finally, we develop new methods to utilize fine-grained knowledge to create generalized and challenging multimodal benchmarks, designed specifically to evaluate future AI models on their multimodal reasoning capabilities.Throughout our research, we conduct extensive experiments to demonstrate the effectiveness of utilizing fine-grained knowledge in improving AI models for multimodal reasoning. Our work focuses on four key perspectives: knowledge extraction, model learning, explainability, and evaluation, providing a comprehensive approach to advancing the field of AI multimodal reasoning.
일반주제명  
Computer science
일반주제명  
Computer engineering
키워드  
Reasoning
키워드  
Machine learning
키워드  
Multimodal learning
키워드  
Multimodal representation
키워드  
Visual entailment
기타저자  
Columbia University Computer Science
기본자료저록  
Dissertations Abstracts International. 86-04B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017164215
■00520250211152934
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798384486329
■035    ▼a(MiAaPQ)AAI31563189
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aWang,  Zhecan.
■24510▼aMultimodal  Reasoning  With  Fine-Grained  Knowledge  Representation
■260    ▼a[Sl]▼bColumbia  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a135  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-04,  Section:  B.
■500    ▼aAdvisor:  Chang,  Shih-Fu.
■5021  ▼aThesis  (Ph.D.)--Columbia  University,  2024.
■520    ▼aMultimodal  reasoning,  especially  the  processing  that  involves  commonsense,  is  a  vital  capability  for  humans,  encompassing  a  wide  range  of  practical  applications,  from  understanding  visual  cues  while  driving  to  interpreting  emotions  and  intentions  in  social  interactions  and  efficiently  planning  and  executing  household  chores.  Therefore,  multimodal  (commonsense)  reasoning  represents  an  important  step  when  developing  advanced  AI  systems  that  aim  to  imitate  human-level  capabilities.  However,  existing  methods  struggle  to  achieve  this  due  to  several  factors,  including  the  under-utilization  of  fine-grained  multimodal  information,  lack  of  transparency,  and  unexplainable  and  unreliable  behaviors  of  the  AI  models.Our  research  addresses  these  challenges  by  focusing  on  improving  AI  models'  multimodal  (commonsense)  reasoning  through  the  utilization  of  fine-grained  knowledge  representation.  We  begin  by  developing  transformer-based  models  to  extract  fine-grained  knowledge  across  various  modalities.  We  then  propose  novel  solutions  to  leverage  this  extracted  knowledge  to  enhance  AI  models'  learning  of  multimodal  reasoning,  particularly  in  downstream  vision-language  understanding  tasks  such  as  visual  question  answering  and  visual  entailment.Beyond  the  focus  on  high  performance,  we  further  propose  approaches  that  exploit  fine-grained  multimodal  knowledge  to  enhance  our  understanding  of  AI  models,  thereby  improving  the  explainability  of  how  vision-language  models  work  during  multimodal  (commonsense)  reasoning.  Finally,  we  develop  new  methods  to  utilize  fine-grained  knowledge  to  create  generalized  and  challenging  multimodal  benchmarks,  designed  specifically  to  evaluate  future  AI  models  on  their  multimodal  reasoning  capabilities.Throughout  our  research,  we  conduct  extensive  experiments  to  demonstrate  the  effectiveness  of  utilizing  fine-grained  knowledge  in  improving  AI  models  for  multimodal  reasoning.  Our  work  focuses  on  four  key  perspectives:  knowledge  extraction,  model  learning,  explainability,  and  evaluation,  providing  a  comprehensive  approach  to  advancing  the  field  of  AI  multimodal  reasoning.
■590    ▼aSchool  code:  0054.
■650  4▼aComputer  science
■650  4▼aComputer  engineering
■653    ▼aReasoning
■653    ▼aMachine  learning
■653    ▼aMultimodal  learning
■653    ▼aMultimodal  representation
■653    ▼aVisual  entailment
■690    ▼a0800
■690    ▼a0984
■690    ▼a0464
■71020▼aColumbia  University▼bComputer  Science.
■7730  ▼tDissertations  Abstracts  International▼g86-04B.
■790    ▼a0054
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164215▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF10827 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.