본문

서브메뉴

Video Models of People and Pixels
Video Models of People and Pixels
Video Models of People and Pixels

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202103601
ISBN  
9798291598498
DDC  
620
저자명  
Rajasegaran, Jathushan.
서명/저자  
Video Models of People and Pixels
발행사항  
[Sl] : University of California, Berkeley, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
86 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-03, Section: B.
주기사항  
Advisor: Malik, Jitendra.
학위논문주기  
Thesis (Ph.D.)--University of California, Berkeley, 2025.
초록/해제  
요약From the moment we are born, we continuously witness the "video" of our own lives-hundreds of thousands of hours of rich, unfolding scenes. These visual experiences, streaming in seamlessly over time, form the foundation of how we understand the world: by tracking motion, recognizing people, and anticipating what comes next. In many ways, our perception begins with tracking-following a pixel, a person, or a motion-enabling higher-order understanding such as object permanence, social interaction, and physical causality. This thesis explores how to build visual models that can track, recognize, and predict.First I will discuss about tracking people in monocular videos with PHALP (Predicting Human Appearance, Location, and Pose). By aggregating 3D representations into tracklets, temporal models predict future states, enabling persistent tracking. Next, I will discuss human action recognition from a Lagrangian perspective using these tracklets. LART (Lagrangian Action Recognition with Tracking), a transformer-based model, demonstrates the benefits of explicit 3D pose (SMPL) and location for predicting actions. LART fuses 3D pose dynamics with contextualized appearance features along tracklets, significantly improving performance on the AVA dataset, especially for interactive and complex actions. Finally, I will discuss about large-scale self-supervised learning through autoregressive video prediction with Toto, a family of causal transformers. Trained on next-token prediction using over a trillion visual tokens from diverse image and video datasets, Toto learns powerful, general-purpose visual representations with minimal inductive biases. An empirical study of architectural and tokenization choices shows these representations achieve competitive performance on downstream tasks including classification, tracking, object permanence, and robotics. We also analyze the power-law scaling of these video models.
일반주제명  
Engineering
일반주제명  
Computer science
일반주제명  
Electrical engineering
키워드  
Action recognition
키워드  
Self-supervised pretraining
키워드  
Tracking people
키워드  
Video models
기타저자  
University of California, Berkeley Electrical Engineering & Computer Sciences
기본자료저록  
Dissertations Abstracts International. 87-03B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017357797
■00520260202103601
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798291598498
■035    ▼a(MiAaPQ)AAI32042494
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a620
■1001  ▼aRajasegaran,  Jathushan.
■24510▼aVideo  Models  of  People  and  Pixels
■260    ▼a[Sl]▼bUniversity  of  California,  Berkeley▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a86  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-03,  Section:  B.
■500    ▼aAdvisor:  Malik,  Jitendra.
■5021  ▼aThesis  (Ph.D.)--University  of  California,  Berkeley,  2025.
■520    ▼aFrom  the  moment  we  are  born,  we  continuously  witness  the  "video"  of  our  own  lives-hundreds  of  thousands  of  hours  of  rich,  unfolding  scenes.  These  visual  experiences,  streaming  in  seamlessly  over  time,  form  the  foundation  of  how  we  understand  the  world:  by  tracking  motion,  recognizing  people,  and  anticipating  what  comes  next.  In  many  ways,  our  perception  begins  with  tracking-following  a  pixel,  a  person,  or  a  motion-enabling  higher-order  understanding  such  as  object  permanence,  social  interaction,  and  physical  causality.  This  thesis  explores  how  to  build  visual  models  that  can  track,  recognize,  and  predict.First  I  will  discuss  about  tracking  people  in  monocular  videos  with  PHALP  (Predicting  Human  Appearance,  Location,  and  Pose).  By  aggregating  3D  representations  into  tracklets,  temporal  models  predict  future  states,  enabling  persistent  tracking.  Next,  I  will  discuss  human  action  recognition  from  a  Lagrangian  perspective  using  these  tracklets.  LART  (Lagrangian  Action  Recognition  with  Tracking),  a  transformer-based  model,  demonstrates  the  benefits  of  explicit  3D  pose  (SMPL)  and  location  for  predicting  actions.  LART  fuses  3D  pose  dynamics  with  contextualized  appearance  features  along  tracklets,  significantly  improving  performance  on  the  AVA  dataset,  especially  for  interactive  and  complex  actions.  Finally,  I  will  discuss  about  large-scale  self-supervised  learning  through  autoregressive  video  prediction  with  Toto,  a  family  of  causal  transformers.  Trained  on  next-token  prediction  using  over  a  trillion  visual  tokens  from  diverse  image  and  video  datasets,  Toto  learns  powerful,  general-purpose  visual  representations  with  minimal  inductive  biases.  An  empirical  study  of  architectural  and  tokenization  choices  shows  these  representations  achieve  competitive  performance  on  downstream  tasks  including  classification,  tracking,  object  permanence,  and  robotics.  We  also  analyze  the  power-law  scaling  of  these  video  models.
■590    ▼aSchool  code:  0028.
■650  4▼aEngineering
■650  4▼aComputer  science
■650  4▼aElectrical  engineering
■653    ▼aAction  recognition
■653    ▼aSelf-supervised  pretraining
■653    ▼aTracking  people
■653    ▼aVideo  models
■690    ▼a0800
■690    ▼a0984
■690    ▼a0544
■690    ▼a0537
■71020▼aUniversity  of  California,  Berkeley▼bElectrical  Engineering  &  Computer  Sciences.
■7730  ▼tDissertations  Abstracts  International▼g87-03B.
■790    ▼a0028
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357797▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF19426 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.