본문

서브메뉴

Unified Approaches for Multi-Task Vision-Language Interactions
Unified Approaches for Multi-Task Vision-Language Interactions
Unified Approaches for Multi-Task Vision-Language Interactions

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211152731
ISBN  
9798384027157
DDC  
401
저자명  
You, Haoxuan.
서명/저자  
Unified Approaches for Multi-Task Vision-Language Interactions
발행사항  
[Sl] : Columbia University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
123 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-02, Section: A.
주기사항  
Advisor: Chang, Shih-Fu.
학위논문주기  
Thesis (Ph.D.)--Columbia University, 2024.
초록/해제  
요약Vision and Language are two major modalities that humans rely on to perceive the environment and understand the world. Recent advances in Artificial Intelligence (AI) facilitate the development of a variety of vision-language tasks derived from diverse multimodal interactions in daily life, such as image captioning, image-text matching, visual question answering (VQA), text-to-image generation, etc. Despite the remarkable performance, most previous state-of-the-art models are merely specialized for a single vision-language task, which lack generalizability across multiple tasks. Additionally, those specialized models sophisticate the algorithm designs and bring redundancy to model deployment when dealing with complex scenes.In this study, we investigate developing unified approaches capable of solving various vision-language interactions in a multi-task manner. We argue that unified multi-task methods could enjoy several potential advantages: (1) A unified framework for multiple tasks can reduce human efforts in designing different models for different tasks; (2) Reusing and sharing parameters across tasks can improve efficiency; (3) Some tasks may be complementary to other tasks so that multi-tasking can boost the performance; (4) They can deal with the complex tasks that need a joint collaborating of multiple basic tasks and enable new applications.In the first part of this thesis, we explore unified multi-task models with the goal of sharing and reusing as many parameters as possible between different tasks. We started with unifying many vision-language question-answering tasks, such as visual entailment, outside-knowledge VQA, and visual commonsense reasoning, in a simple iterative divide-and-conquer framework. Specifically, it iteratively decomposes the original text question into sub-question, solves each sub-question, and derives the answer to the original question, which can uniformly handle reasoning of various types and semantics levels within one framework. In the next work, we take one step further to unify tasks of image-to-text generation, text-to-image generation, vision-language understanding, and image-text matching all in one single large-scale Transformer-based model. The above two works demonstrate the feasibility, effectiveness and efficiency of sharing the parameters across different tasks in a single model. Nevertheless, they still need to switch between different tasks and can only conduct one task at a time.In the second part of this thesis, we introduce our efforts toward simultaneous multi-task models that can conduct multiple tasks at the same time with a single model. It has additional advantages: the model can learn to perform different tasks or combinations of multiple tasks automatically according to user queries; the joint interaction of tasks can enable new potential applications. We begin with compounding spatial understanding and semantic understanding in a single multimodal Transformer-based model. To enable models to understand and localize local regions, we proposed a hybrid region representation that seamlessly bridges regions with image and text. Coupled with a delicately collected training dataset, our model can perform joint spatial and semantic understanding at the same iteration, and empower a new application: spatial reasoning. Continuing the above project, we further introduce an effective module to encode the high-resolution images, and propose a pre-training method that aligns semantics and spatial understanding in high resolution. Besides, we also couple the Optical Character Recognition (OCR) capability together with spatial understanding in the model and study the techniques to improve the compatibility of various tasks.
일반주제명  
Linguistics
일반주제명  
Computer science
키워드  
Deep learning
키워드  
Multi-task
키워드  
Multimodal
키워드  
Vision-language
키워드  
Visual question answering
기타저자  
Columbia University Computer Science
기본자료저록  
Dissertations Abstracts International. 86-02A.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017163615
■00520250211152731
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798384027157
■035    ▼a(MiAaPQ)AAI31490927
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a401
■1001  ▼aYou,  Haoxuan.
■24510▼aUnified  Approaches  for  Multi-Task  Vision-Language  Interactions
■260    ▼a[Sl]▼bColumbia  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a123  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-02,  Section:  A.
■500    ▼aAdvisor:  Chang,  Shih-Fu.
■5021  ▼aThesis  (Ph.D.)--Columbia  University,  2024.
■520    ▼aVision  and  Language  are  two  major  modalities  that  humans  rely  on  to  perceive  the  environment  and  understand  the  world.  Recent  advances  in  Artificial  Intelligence  (AI)  facilitate  the  development  of  a  variety  of  vision-language  tasks  derived  from  diverse  multimodal  interactions  in  daily  life,  such  as  image  captioning,  image-text  matching,  visual  question  answering  (VQA),  text-to-image  generation,  etc.  Despite  the  remarkable  performance,  most  previous  state-of-the-art  models  are  merely  specialized  for  a  single  vision-language  task,  which  lack  generalizability  across  multiple  tasks.  Additionally,  those  specialized  models  sophisticate  the  algorithm  designs  and  bring  redundancy  to  model  deployment  when  dealing  with  complex  scenes.In  this  study,  we  investigate  developing  unified  approaches  capable  of  solving  various  vision-language  interactions  in  a  multi-task  manner.  We  argue  that  unified  multi-task  methods  could  enjoy  several  potential  advantages:  (1)  A  unified  framework  for  multiple  tasks  can  reduce  human  efforts  in  designing  different  models  for  different  tasks;  (2)  Reusing  and  sharing  parameters  across  tasks  can  improve  efficiency;  (3)  Some  tasks  may  be  complementary  to  other  tasks  so  that  multi-tasking  can  boost  the  performance;  (4)  They  can  deal  with  the  complex  tasks  that  need  a  joint  collaborating  of  multiple  basic  tasks  and  enable  new  applications.In  the  first  part  of  this  thesis,  we  explore  unified  multi-task  models  with  the  goal  of  sharing  and  reusing  as  many  parameters  as  possible  between  different  tasks.  We  started  with  unifying  many  vision-language  question-answering  tasks,  such  as  visual  entailment,  outside-knowledge  VQA,  and  visual  commonsense  reasoning,  in  a  simple  iterative  divide-and-conquer  framework.  Specifically,  it  iteratively  decomposes  the  original  text  question  into  sub-question,  solves  each  sub-question,  and  derives  the  answer  to  the  original  question,  which  can  uniformly  handle  reasoning  of  various  types  and  semantics  levels  within  one  framework.  In  the  next  work,  we  take  one  step  further  to  unify  tasks  of  image-to-text  generation,  text-to-image  generation,  vision-language  understanding,  and  image-text  matching  all  in  one  single  large-scale  Transformer-based  model.  The  above  two  works  demonstrate  the  feasibility,  effectiveness  and  efficiency  of  sharing  the  parameters  across  different  tasks  in  a  single  model.  Nevertheless,  they  still  need  to  switch  between  different  tasks  and  can  only  conduct  one  task  at  a  time.In  the  second  part  of  this  thesis,  we  introduce  our  efforts  toward  simultaneous  multi-task  models  that  can  conduct  multiple  tasks  at  the  same  time  with  a  single  model.  It  has  additional  advantages:  the  model  can  learn  to  perform  different  tasks  or  combinations  of  multiple  tasks  automatically  according  to  user  queries;  the  joint  interaction  of  tasks  can  enable  new  potential  applications.  We  begin  with  compounding  spatial  understanding  and  semantic  understanding  in  a  single  multimodal  Transformer-based  model.  To  enable  models  to  understand  and  localize  local  regions,  we  proposed  a  hybrid  region  representation  that  seamlessly  bridges  regions  with  image  and  text.  Coupled  with  a  delicately  collected  training  dataset,  our  model  can  perform  joint  spatial  and  semantic  understanding  at  the  same  iteration,  and  empower  a  new  application:  spatial  reasoning.  Continuing  the  above  project,  we  further  introduce  an  effective  module  to  encode  the  high-resolution  images,  and  propose  a  pre-training  method  that  aligns  semantics  and  spatial  understanding  in  high  resolution.  Besides,  we  also  couple  the  Optical  Character  Recognition  (OCR)  capability  together  with  spatial  understanding  in  the  model  and  study  the  techniques  to  improve  the  compatibility  of  various  tasks.
■590    ▼aSchool  code:  0054.
■650  4▼aLinguistics
■650  4▼aComputer  science
■653    ▼aDeep  learning
■653    ▼aMulti-task
■653    ▼aMultimodal
■653    ▼aVision-language
■653    ▼aVisual  question  answering
■690    ▼a0800
■690    ▼a0984
■690    ▼a0290
■71020▼aColumbia  University▼bComputer  Science.
■7730  ▼tDissertations  Abstracts  International▼g86-02A.
■790    ▼a0054
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17163615▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF13217 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.