서브메뉴
검색
Video Models of People and Pixels
Video Models of People and Pixels
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103601
- ISBN
- 9798291598498
- DDC
- 620
- 서명/저자
- Video Models of People and Pixels
- 발행사항
- [Sl] : University of California, Berkeley, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 86 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-03, Section: B.
- 주기사항
- Advisor: Malik, Jitendra.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Berkeley, 2025.
- 초록/해제
- 요약From the moment we are born, we continuously witness the "video" of our own lives-hundreds of thousands of hours of rich, unfolding scenes. These visual experiences, streaming in seamlessly over time, form the foundation of how we understand the world: by tracking motion, recognizing people, and anticipating what comes next. In many ways, our perception begins with tracking-following a pixel, a person, or a motion-enabling higher-order understanding such as object permanence, social interaction, and physical causality. This thesis explores how to build visual models that can track, recognize, and predict.First I will discuss about tracking people in monocular videos with PHALP (Predicting Human Appearance, Location, and Pose). By aggregating 3D representations into tracklets, temporal models predict future states, enabling persistent tracking. Next, I will discuss human action recognition from a Lagrangian perspective using these tracklets. LART (Lagrangian Action Recognition with Tracking), a transformer-based model, demonstrates the benefits of explicit 3D pose (SMPL) and location for predicting actions. LART fuses 3D pose dynamics with contextualized appearance features along tracklets, significantly improving performance on the AVA dataset, especially for interactive and complex actions. Finally, I will discuss about large-scale self-supervised learning through autoregressive video prediction with Toto, a family of causal transformers. Trained on next-token prediction using over a trillion visual tokens from diverse image and video datasets, Toto learns powerful, general-purpose visual representations with minimal inductive biases. An empirical study of architectural and tokenization choices shows these representations achieve competitive performance on downstream tasks including classification, tracking, object permanence, and robotics. We also analyze the power-law scaling of these video models.
- 일반주제명
- Engineering
- 일반주제명
- Computer science
- 일반주제명
- Electrical engineering
- 키워드
- Tracking people
- 키워드
- Video models
- 기타저자
- University of California, Berkeley Electrical Engineering & Computer Sciences
- 기본자료저록
- Dissertations Abstracts International. 87-03B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017357797
■00520260202103601
■006m o d
■007cr#unu||||||||
■020 ▼a9798291598498
■035 ▼a(MiAaPQ)AAI32042494
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a620
■1001 ▼aRajasegaran, Jathushan.
■24510▼aVideo Models of People and Pixels
■260 ▼a[Sl]▼bUniversity of California, Berkeley▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a86 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-03, Section: B.
■500 ▼aAdvisor: Malik, Jitendra.
■5021 ▼aThesis (Ph.D.)--University of California, Berkeley, 2025.
■520 ▼aFrom the moment we are born, we continuously witness the "video" of our own lives-hundreds of thousands of hours of rich, unfolding scenes. These visual experiences, streaming in seamlessly over time, form the foundation of how we understand the world: by tracking motion, recognizing people, and anticipating what comes next. In many ways, our perception begins with tracking-following a pixel, a person, or a motion-enabling higher-order understanding such as object permanence, social interaction, and physical causality. This thesis explores how to build visual models that can track, recognize, and predict.First I will discuss about tracking people in monocular videos with PHALP (Predicting Human Appearance, Location, and Pose). By aggregating 3D representations into tracklets, temporal models predict future states, enabling persistent tracking. Next, I will discuss human action recognition from a Lagrangian perspective using these tracklets. LART (Lagrangian Action Recognition with Tracking), a transformer-based model, demonstrates the benefits of explicit 3D pose (SMPL) and location for predicting actions. LART fuses 3D pose dynamics with contextualized appearance features along tracklets, significantly improving performance on the AVA dataset, especially for interactive and complex actions. Finally, I will discuss about large-scale self-supervised learning through autoregressive video prediction with Toto, a family of causal transformers. Trained on next-token prediction using over a trillion visual tokens from diverse image and video datasets, Toto learns powerful, general-purpose visual representations with minimal inductive biases. An empirical study of architectural and tokenization choices shows these representations achieve competitive performance on downstream tasks including classification, tracking, object permanence, and robotics. We also analyze the power-law scaling of these video models.
■590 ▼aSchool code: 0028.
■650 4▼aEngineering
■650 4▼aComputer science
■650 4▼aElectrical engineering
■653 ▼aAction recognition
■653 ▼aSelf-supervised pretraining
■653 ▼aTracking people
■653 ▼aVideo models
■690 ▼a0800
■690 ▼a0984
■690 ▼a0544
■690 ▼a0537
■71020▼aUniversity of California, Berkeley▼bElectrical Engineering & Computer Sciences.
■7730 ▼tDissertations Abstracts International▼g87-03B.
■790 ▼a0028
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357797▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


