서브메뉴
검색
Multimodal Representations for Video
Multimodal Representations for Video
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211151137
- ISBN
- 9798382767857
- DDC
- 004
- 서명/저자
- Multimodal Representations for Video
- 발행사항
- [Sl] : Columbia University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 304 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 85-11, Section: B.
- 주기사항
- Advisor: Vondrick, Carl M.
- 학위논문주기
- Thesis (Ph.D.)--Columbia University, 2024.
- 초록/해제
- 요약My thesis explores the fields of multimodal and video analysis in computer vision, aiming to bridge the gap between human perception and machine understanding. Recognizing the interplay among various signals such as text, audio, and visual data, my research explores novel frameworks to integrate these diverse modalities in order to achieve a deeper understanding of complex scenes, with a particular emphasis on video analysis. As part of this exploration, I study diverse tasks such as translation, future prediction, or visual question answering, all connected through the lens of multimodal and video representations. I present novel approaches for each of these challenges, contributing across different facets of computer vision, from dataset creation to algorithmic innovations, and from achieving state-of-the-art results on established benchmarks to introducing new tasks.Methodologically, my thesis embraces two key approaches: self-supervised learning and the integration of structured representations. Self-supervised learning, a technique that allows computers to learn from unlabeled data, helps uncovering inherent connections within multimodal and video inputs. Structured representations, on the other hand, serve as a means to capture complex temporal patterns and uncertainties inherent in video analysis. By employing these techniques, I offer novel insights into modeling multimodal representations for video analysis, showing improved performance with respect to prior work in all studied scenarios.
- 일반주제명
- Computer science
- 키워드
- Computer vision
- 키워드
- Video analysis
- 기타저자
- Columbia University Computer Science
- 기본자료저록
- Dissertations Abstracts International. 85-11B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017160922
■00520250211151137
■006m o d
■007cr#unu||||||||
■020 ▼a9798382767857
■035 ▼a(MiAaPQ)AAI31148554
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aSuris Coll-Vinent, Didac.
■24510▼aMultimodal Representations for Video
■260 ▼a[Sl]▼bColumbia University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a304 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 85-11, Section: B.
■500 ▼aAdvisor: Vondrick, Carl M.
■5021 ▼aThesis (Ph.D.)--Columbia University, 2024.
■520 ▼aMy thesis explores the fields of multimodal and video analysis in computer vision, aiming to bridge the gap between human perception and machine understanding. Recognizing the interplay among various signals such as text, audio, and visual data, my research explores novel frameworks to integrate these diverse modalities in order to achieve a deeper understanding of complex scenes, with a particular emphasis on video analysis. As part of this exploration, I study diverse tasks such as translation, future prediction, or visual question answering, all connected through the lens of multimodal and video representations. I present novel approaches for each of these challenges, contributing across different facets of computer vision, from dataset creation to algorithmic innovations, and from achieving state-of-the-art results on established benchmarks to introducing new tasks.Methodologically, my thesis embraces two key approaches: self-supervised learning and the integration of structured representations. Self-supervised learning, a technique that allows computers to learn from unlabeled data, helps uncovering inherent connections within multimodal and video inputs. Structured representations, on the other hand, serve as a means to capture complex temporal patterns and uncertainties inherent in video analysis. By employing these techniques, I offer novel insights into modeling multimodal representations for video analysis, showing improved performance with respect to prior work in all studied scenarios.
■590 ▼aSchool code: 0054.
■650 4▼aComputer science
■653 ▼aComputer vision
■653 ▼aMultimodal representations
■653 ▼aVideo analysis
■653 ▼aAlgorithmic innovations
■690 ▼a0800
■690 ▼a0984
■71020▼aColumbia University▼bComputer Science.
■7730 ▼tDissertations Abstracts International▼g85-11B.
■790 ▼a0054
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17160922▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


