본문

서브메뉴

Grounding Language in Images and Videos
Grounding Language in Images and Videos
Grounding Language in Images and Videos

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211151509
ISBN  
9798382652443
DDC  
004
저자명  
Sadhu, Arka.
서명/저자  
Grounding Language in Images and Videos
발행사항  
[Sl] : University of Southern California, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
244 p
주기사항  
Source: Dissertations Abstracts International, Volume: 85-11, Section: A.
주기사항  
Advisor: Nevatia, Ramakant.
학위논문주기  
Thesis (Ph.D.)--University of Southern California, 2024.
초록/해제  
요약While machine learning research has traditionally explored image, video and text understanding as separate fields, the surge in multi-modal content in today's digital landscape underscores the importance of computation models that adeptly navigate complex interactions between text, images and videos. This dissertation addresses this challenge of grounding language in visual media - the task of associating linguistic symbols with perceptual experiences and actions. The overarching goal of this dissertation is to bridge the gap between language and vision as a means to a "deeper understanding" of images and videos to allow developing models capable of reasoning over longer-time horizons such as hour-long movies, or a collection of images, or even multiple videos.A pivotal contribution of my work is the use of Semantic Roles for images, videos and text. Unlike previous works that primarily focused on recognizing single entities or generating holistic captions, the use of Semantic Roles facilitates a fine-grained understanding of "who did what to whom" in a structured format. It maintains the advantages of having free-form language phrases and at the same time also being comprehensive and complete like entity recognition, thus enriching the model's interpretive capabilities.In this thesis, we will introduce the various vision-language tasks developed during my Ph.D. This includes grounding unseen words, spatio-temporal localization of entities in a video, video question answering, visual semantic role labeling in videos, reasoning across more than one image or a video, and finally, weakly-supervised open-vocabulary object detection. Each task is accompanied by the creation and development of dedicated datasets, evaluation protocols, and model frameworks. These tasks aim to investigate a particular phenomenon inherent in image or video understanding in isolation, develop corresponding datasets and model frameworks, and outline evaluation protocols robust to data priors.The resulting models can be used for other downstream tasks like obtaining common-sense knowledge graphs from instructional videos or drive end-user applications like Retrieval, Question Answering, and Captioning. By facilitating the deeper integration of language and vision, this dissertation represents a step-forward in machine learning models capable of finer-understanding of the world around us. 
일반주제명  
Computer science
일반주제명  
Computer engineering
일반주제명  
Linguistics
키워드  
Computer vision
키워드  
Image understanding
키워드  
Machine learning
키워드  
Natural language processing
키워드  
Video understanding
기타저자  
University of Southern California Computer Science
기본자료저록  
Dissertations Abstracts International. 85-11A.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017161970
■00520250211151509
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798382652443
■035    ▼a(MiAaPQ)AAI31299441
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aSadhu,  Arka.
■24510▼aGrounding  Language  in  Images  and  Videos
■260    ▼a[Sl]▼bUniversity  of  Southern  California▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a244  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  85-11,  Section:  A.
■500    ▼aAdvisor:  Nevatia,  Ramakant.
■5021  ▼aThesis  (Ph.D.)--University  of  Southern  California,  2024.
■520    ▼aWhile  machine  learning  research  has  traditionally  explored  image,  video  and  text  understanding  as  separate  fields,  the  surge  in  multi-modal  content  in  today's  digital  landscape  underscores  the  importance  of  computation  models  that  adeptly  navigate  complex  interactions  between  text,  images  and  videos.  This  dissertation  addresses  this  challenge  of  grounding  language  in  visual  media  -  the  task  of  associating  linguistic  symbols  with  perceptual  experiences  and  actions.  The  overarching  goal  of  this  dissertation  is  to  bridge  the  gap  between  language  and  vision  as  a  means  to  a  "deeper  understanding"  of  images  and  videos  to  allow  developing  models  capable  of  reasoning  over  longer-time  horizons  such  as  hour-long  movies,  or  a  collection  of  images,  or  even  multiple  videos.A  pivotal  contribution  of  my  work  is  the  use  of  Semantic  Roles  for  images,  videos  and  text.  Unlike  previous  works  that  primarily  focused  on  recognizing  single  entities  or  generating  holistic  captions,  the  use  of  Semantic  Roles  facilitates  a  fine-grained  understanding  of  "who  did  what  to  whom"  in  a  structured  format.  It  maintains  the  advantages  of  having  free-form  language  phrases  and  at  the  same  time  also  being  comprehensive  and  complete  like  entity  recognition,  thus  enriching  the  model's  interpretive  capabilities.In  this  thesis,  we  will  introduce  the  various  vision-language  tasks  developed  during  my  Ph.D.  This  includes  grounding  unseen  words,  spatio-temporal  localization  of  entities  in  a  video,  video  question  answering,  visual  semantic  role  labeling  in  videos,  reasoning  across  more  than  one  image  or  a  video,  and  finally,  weakly-supervised  open-vocabulary  object  detection.  Each  task  is  accompanied  by  the  creation  and  development  of  dedicated  datasets,  evaluation  protocols,  and  model  frameworks.  These  tasks  aim  to  investigate  a  particular  phenomenon  inherent  in  image  or  video  understanding  in  isolation,  develop  corresponding  datasets  and  model  frameworks,  and  outline  evaluation  protocols  robust  to  data  priors.The  resulting  models  can  be  used  for  other  downstream  tasks  like  obtaining  common-sense  knowledge  graphs  from  instructional  videos  or  drive  end-user  applications  like  Retrieval,  Question  Answering,  and  Captioning.  By  facilitating  the  deeper  integration  of  language  and  vision,  this  dissertation  represents  a  step-forward  in  machine  learning  models  capable  of  finer-understanding  of  the  world  around  us. 
■590    ▼aSchool  code:  0208.
■650  4▼aComputer  science
■650  4▼aComputer  engineering
■650  4▼aLinguistics
■653    ▼aComputer  vision
■653    ▼aImage  understanding
■653    ▼aMachine  learning
■653    ▼aNatural  language  processing
■653    ▼aVideo  understanding
■690    ▼a0984
■690    ▼a0464
■690    ▼a0800
■690    ▼a0290
■71020▼aUniversity  of  Southern  California▼bComputer  Science.
■7730  ▼tDissertations  Abstracts  International▼g85-11A.
■790    ▼a0208
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17161970▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF11208 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.