서브메뉴
검색
Towards Generalist Vision-Language Models for Videos in Embodied AI
Towards Generalist Vision-Language Models for Videos in Embodied AI
Detailed Information
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105222
- ISBN
- 9798291566343
- DDC
- 004
- 서명/저자
- Towards Generalist Vision-Language Models for Videos in Embodied AI
- 발행사항
- [Sl] : University of Michigan, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 142 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-03, Section: A.
- 주기사항
- Advisor: Chai, Joyce Y.
- 학위논문주기
- Thesis (Ph.D.)--University of Michigan, 2025.
- 초록/해제
- 요약Recent advances in vision-language models (VLMs)--the multimodal descendants of large language models (LLMs)--have shown tremendous promise in addressing the challenges of embodied AI, particularly in open-world settings. However, existing VLMs are primarily designed for static, turn-based settings and struggle in dynamic, real-time environments with continuous inputs. In this thesis, we address these limitations by investigating three core challenges in adapting VLMs to open-world, real-time embodied AI: (1) addressing the long-tail problem of the multimodal open world, (2) processing long videos spanning minutes to hours, and (3) enabling real-time interaction in continuously changing environments.To address the first challenge, we identify in-context learning as key to tackling the long-tail problem of the multimodal open world, and propose Emergent In-context Learning on Videos (EILeV), a novel training paradigm that induces in-context learning capabilities in VLMs over video and text. EILeV achieves this by curating a training dataset with specific distributional properties and training a VLM with architectural modifications that enable it to process inputs interleaved with video and text.To address the second challenge, we introduce Espresso, a novel projector architecture that encodes long videos using a fixed number of tokens without sacrificing the VLM's temporal understanding. Espresso achieves this by separately compressing spatial and temporal features, making efficient use of the fixed token budget when encoding video inputs.Finally, to address the third challenge, we introduce Temporally-Grounded Language Generation (TGLG), a benchmark task that evaluates two critical capabilities for real-time VLMs: perceptual updating--the ability to account for environmental changes while generating a response, and contingency awareness--the ability to adjust responses based on how previous outputs affect the environment. We curate a video-text dataset for this task and propose Temporal Responsiveness and Alignment Coherence Evaluation (TRACE), a new metric for quantifying these capabilities. As a strong baseline, we present Vision-Language Models with Time-Synchronized Interleaving (VLM-TSI), which tightly interleaves vision and text tokens to model real-time interactions with high temporal fidelity.By addressing these three challenges, this thesis advances the development of VLMs that can reason and respond fluidly in open-world, real-time environments.
- 일반주제명
- Computer science
- 일반주제명
- Computer engineering
- 일반주제명
- Information science
- 기타저자
- University of Michigan Computer Science & Engineering
- 기본자료저록
- Dissertations Abstracts International. 87-03A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017359838
■00520260202105222
■006m o d
■007cr#unu||||||||
■020 ▼a9798291566343
■035 ▼a(MiAaPQ)AAI32271819
■035 ▼a(MiAaPQ)umichrackham006371
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aYu, Keunwoo Peter.
■24510▼aTowards Generalist Vision-Language Models for Videos in Embodied AI
■260 ▼a[Sl]▼bUniversity of Michigan▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a142 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-03, Section: A.
■500 ▼aAdvisor: Chai, Joyce Y.
■5021 ▼aThesis (Ph.D.)--University of Michigan, 2025.
■520 ▼aRecent advances in vision-language models (VLMs)--the multimodal descendants of large language models (LLMs)--have shown tremendous promise in addressing the challenges of embodied AI, particularly in open-world settings. However, existing VLMs are primarily designed for static, turn-based settings and struggle in dynamic, real-time environments with continuous inputs. In this thesis, we address these limitations by investigating three core challenges in adapting VLMs to open-world, real-time embodied AI: (1) addressing the long-tail problem of the multimodal open world, (2) processing long videos spanning minutes to hours, and (3) enabling real-time interaction in continuously changing environments.To address the first challenge, we identify in-context learning as key to tackling the long-tail problem of the multimodal open world, and propose Emergent In-context Learning on Videos (EILeV), a novel training paradigm that induces in-context learning capabilities in VLMs over video and text. EILeV achieves this by curating a training dataset with specific distributional properties and training a VLM with architectural modifications that enable it to process inputs interleaved with video and text.To address the second challenge, we introduce Espresso, a novel projector architecture that encodes long videos using a fixed number of tokens without sacrificing the VLM's temporal understanding. Espresso achieves this by separately compressing spatial and temporal features, making efficient use of the fixed token budget when encoding video inputs.Finally, to address the third challenge, we introduce Temporally-Grounded Language Generation (TGLG), a benchmark task that evaluates two critical capabilities for real-time VLMs: perceptual updating--the ability to account for environmental changes while generating a response, and contingency awareness--the ability to adjust responses based on how previous outputs affect the environment. We curate a video-text dataset for this task and propose Temporal Responsiveness and Alignment Coherence Evaluation (TRACE), a new metric for quantifying these capabilities. As a strong baseline, we present Vision-Language Models with Time-Synchronized Interleaving (VLM-TSI), which tightly interleaves vision and text tokens to model real-time interactions with high temporal fidelity.By addressing these three challenges, this thesis advances the development of VLMs that can reason and respond fluidly in open-world, real-time environments.
■590 ▼aSchool code: 0127.
■650 4▼aComputer science
■650 4▼aComputer engineering
■650 4▼aInformation science
■653 ▼aVision-language models
■653 ▼aLarge language models
■653 ▼aTemporally-Grounded Language Generation
■653 ▼aEmergent In-context Learning on Videos
■690 ▼a0984
■690 ▼a0464
■690 ▼a0800
■690 ▼a0723
■71020▼aUniversity of Michigan▼bComputer Science & Engineering.
■7730 ▼tDissertations Abstracts International▼g87-03A.
■790 ▼a0127
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17359838▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.
Preview
Export
ChatGPT Discussion
AI Recommended Related Books
Подробнее информация.
- Бронирование
- не существует
- моя папка
- Первый запрос зрения
- Non-Book Loan Application
- Nighttime Book Loan Application
Available after logging in.


