본문

서브메뉴

Towards Video Understanding Through Language in Real-Life Settings
Towards Video Understanding Through Language in Real-Life Settings
Towards Video Understanding Through Language in Real-Life Settings

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211153014
ISBN  
9798384045496
DDC  
004
저자명  
Castro, Santiago.
서명/저자  
Towards Video Understanding Through Language in Real-Life Settings
발행사항  
[Sl] : University of Michigan, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
193 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-04, Section: A.
주기사항  
Advisor: Mihalcea, Rada.
학위논문주기  
Thesis (Ph.D.)--University of Michigan, 2024.
초록/해제  
요약Videos have become an integral part of our daily lives, with a rapidly growing number on YouTube, Netflix, and TikTok serving as testimony to their widespread popularity. Behind the simplicity of their interfaces and user experiences, the systems that power these products employ numerous video-understanding techniques, even for straightforward use cases such as finding a video on how to cook salmon. Despite the significant progress achieved in this area, there remains a gap between lab-setting capabilities and reality, as multiple phenomena are not adequately designed for realistic settings, causing various issues such as domain mismatches and the diverse way people interact in videos (e.g., sarcastically). My work aims to bridge this gap by enabling the understanding of video content in realistic settings.The issues that make current video understanding research unsuitable for real life can be classified into data, methods, and evaluation. The data aspect is crucial since current research has predominantly overlooked real-life settings. I present new datasets and benchmarks for such domains: daily situations and in-the-wild scenarios. These benchmarks measure the effectiveness of new methods in these more realistic settings. Likewise, I introduce a novel framework that accounts for a typical yet understudied human behavior: sarcasm. Sarcasm is particularly suited to be studied in video since I show that leveraging what we see and hear (as people commonly do) allows one to understand it better. For the methods aspect, I consider a fundamental issue, which is the impracticality and lack of scalability of the traditional in-the-lab setting, tuning one model for each newly addressed task and domain. I propose a robust method that allows practitioners to employ a single model for novel tasks and domains with satisfactory performance. Additionally, I present a technique to improve the compositional generalization of existing models. Finally, I focus on current practices for evaluation and propose a framework better suited to realistic settings. Current benchmarks for short video understanding have drawbacks, such as employing easy-to-detect distractor answers, not accounting for diversity when depicting the same situation, and not considering realistic settings. I present a novel evaluation format that tackles all these issues and a benchmark that leverages it. The benchmark shows a gap between the performance of several methods and humans.
일반주제명  
Computer science
일반주제명  
Computer engineering
일반주제명  
Web studies
일반주제명  
Information technology
키워드  
Video understanding
키워드  
Natural Language Processing
키워드  
Computer Vision
키워드  
Compositional generalization
키워드  
Sarcasm
기타저자  
University of Michigan Computer Science & Engineering
기본자료저록  
Dissertations Abstracts International. 86-04A.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017164532
■00520250211153014
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798384045496
■035    ▼a(MiAaPQ)AAI31631493
■035    ▼a(MiAaPQ)umichrackham005804
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aCastro,  Santiago.
■24510▼aTowards  Video  Understanding  Through  Language  in  Real-Life  Settings
■260    ▼a[Sl]▼bUniversity  of  Michigan▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a193  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-04,  Section:  A.
■500    ▼aAdvisor:  Mihalcea,  Rada.
■5021  ▼aThesis  (Ph.D.)--University  of  Michigan,  2024.
■520    ▼aVideos  have  become  an  integral  part  of  our  daily  lives,  with  a  rapidly  growing  number  on  YouTube,  Netflix,  and  TikTok  serving  as  testimony  to  their  widespread  popularity.  Behind  the  simplicity  of  their  interfaces  and  user  experiences,  the  systems  that  power  these  products  employ  numerous  video-understanding  techniques,  even  for  straightforward  use  cases  such  as  finding  a  video  on  how  to  cook  salmon.  Despite  the  significant  progress  achieved  in  this  area,  there  remains  a  gap  between  lab-setting  capabilities  and  reality,  as  multiple  phenomena  are  not  adequately  designed  for  realistic  settings,  causing  various  issues  such  as  domain  mismatches  and  the  diverse  way  people  interact  in  videos  (e.g.,  sarcastically).  My  work  aims  to  bridge  this  gap  by  enabling  the  understanding  of  video  content  in  realistic  settings.The  issues  that  make  current  video  understanding  research  unsuitable  for  real  life  can  be  classified  into  data,  methods,  and  evaluation.  The  data  aspect  is  crucial  since  current  research  has  predominantly  overlooked  real-life  settings.  I  present  new  datasets  and  benchmarks  for  such  domains:  daily  situations  and  in-the-wild  scenarios.  These  benchmarks  measure  the  effectiveness  of  new  methods  in  these  more  realistic  settings.  Likewise,  I  introduce  a  novel  framework  that  accounts  for  a  typical  yet  understudied  human  behavior:  sarcasm.  Sarcasm  is  particularly  suited  to  be  studied  in  video  since  I  show  that  leveraging  what  we  see  and  hear  (as  people  commonly  do)  allows  one  to  understand  it  better.  For  the  methods  aspect,  I  consider  a  fundamental  issue,  which  is  the  impracticality  and  lack  of  scalability  of  the  traditional  in-the-lab  setting,  tuning  one  model  for  each  newly  addressed  task  and  domain.  I  propose  a  robust  method  that  allows  practitioners  to  employ  a  single  model  for  novel  tasks  and  domains  with  satisfactory  performance.  Additionally,  I  present  a  technique  to  improve  the  compositional  generalization  of  existing  models.  Finally,  I  focus  on  current  practices  for  evaluation  and  propose  a  framework  better  suited  to  realistic  settings.  Current  benchmarks  for  short  video  understanding  have  drawbacks,  such  as  employing  easy-to-detect  distractor  answers,  not  accounting  for  diversity  when  depicting  the  same  situation,  and  not  considering  realistic  settings.  I  present  a  novel  evaluation  format  that  tackles  all  these  issues  and  a  benchmark  that  leverages  it.  The  benchmark  shows  a  gap  between  the  performance  of  several  methods  and  humans.
■590    ▼aSchool  code:  0127.
■650  4▼aComputer  science
■650  4▼aComputer  engineering
■650  4▼aWeb  studies
■650  4▼aInformation  technology
■653    ▼aVideo  understanding
■653    ▼aNatural  Language  Processing
■653    ▼aComputer  Vision
■653    ▼aCompositional  generalization
■653    ▼aSarcasm  
■690    ▼a0984
■690    ▼a0800
■690    ▼a0489
■690    ▼a0464
■690    ▼a0646
■71020▼aUniversity  of  Michigan▼bComputer  Science  &  Engineering.
■7730  ▼tDissertations  Abstracts  International▼g86-04A.
■790    ▼a0127
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164532▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF11394 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.