본문

서브메뉴

Multimodal Representations for Video
Multimodal Representations for Video
Multimodal Representations for Video

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211151137
ISBN  
9798382767857
DDC  
004
저자명  
Suris Coll-Vinent, Didac.
서명/저자  
Multimodal Representations for Video
발행사항  
[Sl] : Columbia University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
304 p
주기사항  
Source: Dissertations Abstracts International, Volume: 85-11, Section: B.
주기사항  
Advisor: Vondrick, Carl M.
학위논문주기  
Thesis (Ph.D.)--Columbia University, 2024.
초록/해제  
요약My thesis explores the fields of multimodal and video analysis in computer vision, aiming to bridge the gap between human perception and machine understanding. Recognizing the interplay among various signals such as text, audio, and visual data, my research explores novel frameworks to integrate these diverse modalities in order to achieve a deeper understanding of complex scenes, with a particular emphasis on video analysis. As part of this exploration, I study diverse tasks such as translation, future prediction, or visual question answering, all connected through the lens of multimodal and video representations. I present novel approaches for each of these challenges, contributing across different facets of computer vision, from dataset creation to algorithmic innovations, and from achieving state-of-the-art results on established benchmarks to introducing new tasks.Methodologically, my thesis embraces two key approaches: self-supervised learning and the integration of structured representations. Self-supervised learning, a technique that allows computers to learn from unlabeled data, helps uncovering inherent connections within multimodal and video inputs. Structured representations, on the other hand, serve as a means to capture complex temporal patterns and uncertainties inherent in video analysis. By employing these techniques, I offer novel insights into modeling multimodal representations for video analysis, showing improved performance with respect to prior work in all studied scenarios.
일반주제명  
Computer science
키워드  
Computer vision
키워드  
Multimodal representations
키워드  
Video analysis
키워드  
Algorithmic innovations
기타저자  
Columbia University Computer Science
기본자료저록  
Dissertations Abstracts International. 85-11B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017160922
■00520250211151137
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798382767857
■035    ▼a(MiAaPQ)AAI31148554
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aSuris  Coll-Vinent,  Didac.
■24510▼aMultimodal  Representations  for  Video
■260    ▼a[Sl]▼bColumbia  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a304  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  85-11,  Section:  B.
■500    ▼aAdvisor:  Vondrick,  Carl  M.
■5021  ▼aThesis  (Ph.D.)--Columbia  University,  2024.
■520    ▼aMy  thesis  explores  the  fields  of  multimodal  and  video  analysis  in  computer  vision,  aiming  to  bridge  the  gap  between  human  perception  and  machine  understanding.  Recognizing  the  interplay  among  various  signals  such  as  text,  audio,  and  visual  data,  my  research  explores  novel  frameworks  to  integrate  these  diverse  modalities  in  order  to  achieve  a  deeper  understanding  of  complex  scenes,  with  a  particular  emphasis  on  video  analysis.  As  part  of  this  exploration,  I  study  diverse  tasks  such  as  translation,  future  prediction,  or  visual  question  answering,  all  connected  through  the  lens  of  multimodal  and  video  representations.  I  present  novel  approaches  for  each  of  these  challenges,  contributing  across  different  facets  of  computer  vision,  from  dataset  creation  to  algorithmic  innovations,  and  from  achieving  state-of-the-art  results  on  established  benchmarks  to  introducing  new  tasks.Methodologically,  my  thesis  embraces  two  key  approaches:  self-supervised  learning  and  the  integration  of  structured  representations.  Self-supervised  learning,  a  technique  that  allows  computers  to  learn  from  unlabeled  data,  helps  uncovering  inherent  connections  within  multimodal  and  video  inputs.  Structured  representations,  on  the  other  hand,  serve  as  a  means  to  capture  complex  temporal  patterns  and  uncertainties  inherent  in  video  analysis.  By  employing  these  techniques,  I  offer  novel  insights  into  modeling  multimodal  representations  for  video  analysis,    showing  improved  performance  with  respect  to  prior  work  in  all  studied  scenarios.
■590    ▼aSchool  code:  0054.
■650  4▼aComputer  science
■653    ▼aComputer  vision
■653    ▼aMultimodal  representations
■653    ▼aVideo  analysis
■653    ▼aAlgorithmic  innovations
■690    ▼a0800
■690    ▼a0984
■71020▼aColumbia  University▼bComputer  Science.
■7730  ▼tDissertations  Abstracts  International▼g85-11B.
■790    ▼a0054
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17160922▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF12062 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.