본문

서브메뉴

Towards Generalist Vision-Language Models for Videos in Embodied AI
Towards Generalist Vision-Language Models for Videos in Embodied AI
Towards Generalist Vision-Language Models for Videos in Embodied AI

Detailed Information

자료유형  
 학위논문 서양
최종처리일시  
20260202105222
ISBN  
9798291566343
DDC  
004
저자명  
Yu, Keunwoo Peter.
서명/저자  
Towards Generalist Vision-Language Models for Videos in Embodied AI
발행사항  
[Sl] : University of Michigan, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
142 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-03, Section: A.
주기사항  
Advisor: Chai, Joyce Y.
학위논문주기  
Thesis (Ph.D.)--University of Michigan, 2025.
초록/해제  
요약Recent advances in vision-language models (VLMs)--the multimodal descendants of large language models (LLMs)--have shown tremendous promise in addressing the challenges of embodied AI, particularly in open-world settings. However, existing VLMs are primarily designed for static, turn-based settings and struggle in dynamic, real-time environments with continuous inputs. In this thesis, we address these limitations by investigating three core challenges in adapting VLMs to open-world, real-time embodied AI: (1) addressing the long-tail problem of the multimodal open world, (2) processing long videos spanning minutes to hours, and (3) enabling real-time interaction in continuously changing environments.To address the first challenge, we identify in-context learning as key to tackling the long-tail problem of the multimodal open world, and propose Emergent In-context Learning on Videos (EILeV), a novel training paradigm that induces in-context learning capabilities in VLMs over video and text. EILeV achieves this by curating a training dataset with specific distributional properties and training a VLM with architectural modifications that enable it to process inputs interleaved with video and text.To address the second challenge, we introduce Espresso, a novel projector architecture that encodes long videos using a fixed number of tokens without sacrificing the VLM's temporal understanding. Espresso achieves this by separately compressing spatial and temporal features, making efficient use of the fixed token budget when encoding video inputs.Finally, to address the third challenge, we introduce Temporally-Grounded Language Generation (TGLG), a benchmark task that evaluates two critical capabilities for real-time VLMs: perceptual updating--the ability to account for environmental changes while generating a response, and contingency awareness--the ability to adjust responses based on how previous outputs affect the environment. We curate a video-text dataset for this task and propose Temporal Responsiveness and Alignment Coherence Evaluation (TRACE), a new metric for quantifying these capabilities. As a strong baseline, we present Vision-Language Models with Time-Synchronized Interleaving (VLM-TSI), which tightly interleaves vision and text tokens to model real-time interactions with high temporal fidelity.By addressing these three challenges, this thesis advances the development of VLMs that can reason and respond fluidly in open-world, real-time environments.
일반주제명  
Computer science
일반주제명  
Computer engineering
일반주제명  
Information science
키워드  
Vision-language models
키워드  
Large language models
키워드  
Temporally-Grounded Language Generation
키워드  
Emergent In-context Learning on Videos
기타저자  
University of Michigan Computer Science & Engineering
기본자료저록  
Dissertations Abstracts International. 87-03A.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017359838
■00520260202105222
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798291566343
■035    ▼a(MiAaPQ)AAI32271819
■035    ▼a(MiAaPQ)umichrackham006371
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aYu,  Keunwoo  Peter.
■24510▼aTowards  Generalist  Vision-Language  Models  for  Videos  in  Embodied  AI
■260    ▼a[Sl]▼bUniversity  of  Michigan▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a142  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-03,  Section:  A.
■500    ▼aAdvisor:  Chai,  Joyce  Y.
■5021  ▼aThesis  (Ph.D.)--University  of  Michigan,  2025.
■520    ▼aRecent  advances  in  vision-language  models  (VLMs)--the  multimodal  descendants  of  large  language  models  (LLMs)--have  shown  tremendous  promise  in  addressing  the  challenges  of  embodied  AI,  particularly  in  open-world  settings.  However,  existing  VLMs  are  primarily  designed  for  static,  turn-based  settings  and  struggle  in  dynamic,  real-time  environments  with  continuous  inputs.  In  this  thesis,  we  address  these  limitations  by  investigating  three  core  challenges  in  adapting  VLMs  to  open-world,  real-time  embodied  AI:  (1)  addressing  the  long-tail  problem  of  the  multimodal  open  world,  (2)  processing  long  videos  spanning  minutes  to  hours,  and  (3)  enabling  real-time  interaction  in  continuously  changing  environments.To  address  the  first  challenge,  we  identify  in-context  learning  as  key  to  tackling  the  long-tail  problem  of  the  multimodal  open  world,  and  propose  Emergent  In-context  Learning  on  Videos  (EILeV),  a  novel  training  paradigm  that  induces  in-context  learning  capabilities  in  VLMs  over  video  and  text.  EILeV  achieves  this  by  curating  a  training  dataset  with  specific  distributional  properties  and  training  a  VLM  with  architectural  modifications  that  enable  it  to  process  inputs  interleaved  with  video  and  text.To  address  the  second  challenge,  we  introduce  Espresso,  a  novel  projector  architecture  that  encodes  long  videos  using  a  fixed  number  of  tokens  without  sacrificing  the  VLM's  temporal  understanding.  Espresso  achieves  this  by  separately  compressing  spatial  and  temporal  features,  making  efficient  use  of  the  fixed  token  budget  when  encoding  video  inputs.Finally,  to  address  the  third  challenge,  we  introduce  Temporally-Grounded  Language  Generation  (TGLG),  a  benchmark  task  that  evaluates  two  critical  capabilities  for  real-time  VLMs:  perceptual  updating--the  ability  to  account  for  environmental  changes  while  generating  a  response,  and  contingency  awareness--the  ability  to  adjust  responses  based  on  how  previous  outputs  affect  the  environment.  We  curate  a  video-text  dataset  for  this  task  and  propose  Temporal  Responsiveness  and  Alignment  Coherence  Evaluation  (TRACE),  a  new  metric  for  quantifying  these  capabilities.  As  a  strong  baseline,  we  present  Vision-Language  Models  with  Time-Synchronized  Interleaving  (VLM-TSI),  which  tightly  interleaves  vision  and  text  tokens  to  model  real-time  interactions  with  high  temporal  fidelity.By  addressing  these  three  challenges,  this  thesis  advances  the  development  of  VLMs  that  can  reason  and  respond  fluidly  in  open-world,  real-time  environments.
■590    ▼aSchool  code:  0127.
■650  4▼aComputer  science
■650  4▼aComputer  engineering
■650  4▼aInformation  science
■653    ▼aVision-language  models
■653    ▼aLarge  language  models
■653    ▼aTemporally-Grounded  Language  Generation
■653    ▼aEmergent  In-context  Learning  on  Videos
■690    ▼a0984
■690    ▼a0464
■690    ▼a0800
■690    ▼a0723
■71020▼aUniversity  of  Michigan▼bComputer  Science  &  Engineering.
■7730  ▼tDissertations  Abstracts  International▼g87-03A.
■790    ▼a0127
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17359838▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

Preview

Export

ChatGPT Discussion

AI Recommended Related Books


    New Books MORE
    Statistics for the past 3 years. Go to brief

    Подробнее информация.

    • Бронирование
    • не существует
    • моя папка
    • Первый запрос зрения
    • Non-Book Loan Application
    • Nighttime Book Loan Application
    материал
    Reg No. Количество платежных Местоположение статус Ленд информации
    TF17921 전자도서 대출가능 My Folder 부재도서신고 비도서대출신청 야간 도서대출신청

    * Бронирование доступны в заимствований книги. Чтобы сделать предварительный заказ, пожалуйста, нажмите кнопку бронирование

    Books borrowed together with this book

    Related Popular Books

    Available after logging in.