서브메뉴
검색
Towards Video Understanding Through Language in Real-Life Settings
Towards Video Understanding Through Language in Real-Life Settings
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211153014
- ISBN
- 9798384045496
- DDC
- 004
- 서명/저자
- Towards Video Understanding Through Language in Real-Life Settings
- 발행사항
- [Sl] : University of Michigan, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 193 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-04, Section: A.
- 주기사항
- Advisor: Mihalcea, Rada.
- 학위논문주기
- Thesis (Ph.D.)--University of Michigan, 2024.
- 초록/해제
- 요약Videos have become an integral part of our daily lives, with a rapidly growing number on YouTube, Netflix, and TikTok serving as testimony to their widespread popularity. Behind the simplicity of their interfaces and user experiences, the systems that power these products employ numerous video-understanding techniques, even for straightforward use cases such as finding a video on how to cook salmon. Despite the significant progress achieved in this area, there remains a gap between lab-setting capabilities and reality, as multiple phenomena are not adequately designed for realistic settings, causing various issues such as domain mismatches and the diverse way people interact in videos (e.g., sarcastically). My work aims to bridge this gap by enabling the understanding of video content in realistic settings.The issues that make current video understanding research unsuitable for real life can be classified into data, methods, and evaluation. The data aspect is crucial since current research has predominantly overlooked real-life settings. I present new datasets and benchmarks for such domains: daily situations and in-the-wild scenarios. These benchmarks measure the effectiveness of new methods in these more realistic settings. Likewise, I introduce a novel framework that accounts for a typical yet understudied human behavior: sarcasm. Sarcasm is particularly suited to be studied in video since I show that leveraging what we see and hear (as people commonly do) allows one to understand it better. For the methods aspect, I consider a fundamental issue, which is the impracticality and lack of scalability of the traditional in-the-lab setting, tuning one model for each newly addressed task and domain. I propose a robust method that allows practitioners to employ a single model for novel tasks and domains with satisfactory performance. Additionally, I present a technique to improve the compositional generalization of existing models. Finally, I focus on current practices for evaluation and propose a framework better suited to realistic settings. Current benchmarks for short video understanding have drawbacks, such as employing easy-to-detect distractor answers, not accounting for diversity when depicting the same situation, and not considering realistic settings. I present a novel evaluation format that tackles all these issues and a benchmark that leverages it. The benchmark shows a gap between the performance of several methods and humans.
- 일반주제명
- Computer science
- 일반주제명
- Computer engineering
- 일반주제명
- Web studies
- 일반주제명
- Information technology
- 키워드
- Computer Vision
- 키워드
- Sarcasm
- 기타저자
- University of Michigan Computer Science & Engineering
- 기본자료저록
- Dissertations Abstracts International. 86-04A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017164532
■00520250211153014
■006m o d
■007cr#unu||||||||
■020 ▼a9798384045496
■035 ▼a(MiAaPQ)AAI31631493
■035 ▼a(MiAaPQ)umichrackham005804
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aCastro, Santiago.
■24510▼aTowards Video Understanding Through Language in Real-Life Settings
■260 ▼a[Sl]▼bUniversity of Michigan▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a193 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-04, Section: A.
■500 ▼aAdvisor: Mihalcea, Rada.
■5021 ▼aThesis (Ph.D.)--University of Michigan, 2024.
■520 ▼aVideos have become an integral part of our daily lives, with a rapidly growing number on YouTube, Netflix, and TikTok serving as testimony to their widespread popularity. Behind the simplicity of their interfaces and user experiences, the systems that power these products employ numerous video-understanding techniques, even for straightforward use cases such as finding a video on how to cook salmon. Despite the significant progress achieved in this area, there remains a gap between lab-setting capabilities and reality, as multiple phenomena are not adequately designed for realistic settings, causing various issues such as domain mismatches and the diverse way people interact in videos (e.g., sarcastically). My work aims to bridge this gap by enabling the understanding of video content in realistic settings.The issues that make current video understanding research unsuitable for real life can be classified into data, methods, and evaluation. The data aspect is crucial since current research has predominantly overlooked real-life settings. I present new datasets and benchmarks for such domains: daily situations and in-the-wild scenarios. These benchmarks measure the effectiveness of new methods in these more realistic settings. Likewise, I introduce a novel framework that accounts for a typical yet understudied human behavior: sarcasm. Sarcasm is particularly suited to be studied in video since I show that leveraging what we see and hear (as people commonly do) allows one to understand it better. For the methods aspect, I consider a fundamental issue, which is the impracticality and lack of scalability of the traditional in-the-lab setting, tuning one model for each newly addressed task and domain. I propose a robust method that allows practitioners to employ a single model for novel tasks and domains with satisfactory performance. Additionally, I present a technique to improve the compositional generalization of existing models. Finally, I focus on current practices for evaluation and propose a framework better suited to realistic settings. Current benchmarks for short video understanding have drawbacks, such as employing easy-to-detect distractor answers, not accounting for diversity when depicting the same situation, and not considering realistic settings. I present a novel evaluation format that tackles all these issues and a benchmark that leverages it. The benchmark shows a gap between the performance of several methods and humans.
■590 ▼aSchool code: 0127.
■650 4▼aComputer science
■650 4▼aComputer engineering
■650 4▼aWeb studies
■650 4▼aInformation technology
■653 ▼aVideo understanding
■653 ▼aNatural Language Processing
■653 ▼aComputer Vision
■653 ▼aCompositional generalization
■653 ▼aSarcasm
■690 ▼a0984
■690 ▼a0800
■690 ▼a0489
■690 ▼a0464
■690 ▼a0646
■71020▼aUniversity of Michigan▼bComputer Science & Engineering.
■7730 ▼tDissertations Abstracts International▼g86-04A.
■790 ▼a0127
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164532▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


