본문

서브메뉴

Generative Models of Vision and Action
Generative Models of Vision and Action
Generative Models of Vision and Action

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211152748
ISBN  
9798342107396
DDC  
620
저자명  
Gupta, Agrim.
서명/저자  
Generative Models of Vision and Action
발행사항  
[Sl] : Stanford University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
129 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-04, Section: B.
주기사항  
Advisor: Li, Fei-Fei.
학위논문주기  
Thesis (Ph.D.)--Stanford University, 2024.
초록/해제  
요약Animals and humans display remarkable ability at building internal representations of the world and using them to simulate, evaluate and select among different possible actions. This capability is learnt primarily from observation and without any supervision. Endowing autonomous agents with similar capabilities is a fundamental challenge in machine learning. In this thesis I will explore new algorithms that enable scalable representation learning from videos via prediction, generative models of visual data and their applications to robotics.To begin, I will discuss the challenges associated with using predictive learning objectives to learn visual representations. I'll introduce a simple predictive learning architecture and objective that enables learning visual representations capable of solving a wide range of visual correspondence tasks in a zero-shot manner. Subsequently, I'll present a transformer-based approach for photorealistic video generation via diffusion modeling. Our approach jointly compresses images and videos within a unified latent space, enabling training and generation across modalities. Finally, I will illustrate the practical applications of generative models for robot learning. Our non-autoregressive, action-conditioned video generation model can act as a world model, enabling embodied agents to plan using visual model-predictive control. Furthermore, I'll showcase a generalist agent trained via next token prediction to learn from diverse robotic experiences across various robots and tasks.
일반주제명  
Robots
일반주제명  
Success
일반주제명  
Failure analysis
일반주제명  
Video recordings
일반주제명  
Semantics
일반주제명  
Film studies
일반주제명  
Logic
일반주제명  
Robotics
기타저자  
Stanford University.
기본자료저록  
Dissertations Abstracts International. 86-04B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017163752
■00520250211152748
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798342107396
■035    ▼a(MiAaPQ)AAI31520329
■035    ▼a(MiAaPQ)Stanfordwd022wx6061
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a620
■1001  ▼aGupta,  Agrim.
■24510▼aGenerative  Models  of  Vision  and  Action
■260    ▼a[Sl]▼bStanford  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a129  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-04,  Section:  B.
■500    ▼aAdvisor:  Li,  Fei-Fei.
■5021  ▼aThesis  (Ph.D.)--Stanford  University,  2024.
■520    ▼aAnimals  and  humans  display  remarkable  ability  at  building  internal  representations  of  the  world  and  using  them  to  simulate,  evaluate  and  select  among  different  possible  actions.  This  capability  is  learnt  primarily  from  observation  and  without  any  supervision.  Endowing  autonomous  agents  with  similar  capabilities  is  a  fundamental  challenge  in  machine  learning.  In  this  thesis  I  will  explore  new  algorithms  that  enable  scalable  representation  learning  from  videos  via  prediction,  generative  models  of  visual  data  and  their  applications  to  robotics.To  begin,  I  will  discuss  the  challenges  associated  with  using  predictive  learning  objectives  to  learn  visual  representations.  I'll  introduce  a  simple  predictive  learning  architecture  and  objective  that  enables  learning  visual  representations  capable  of  solving  a  wide  range  of  visual  correspondence  tasks  in  a  zero-shot  manner.  Subsequently,  I'll  present  a  transformer-based  approach  for  photorealistic  video  generation  via  diffusion  modeling.  Our  approach  jointly  compresses  images  and  videos  within  a  unified  latent  space,  enabling  training  and  generation  across  modalities.  Finally,  I  will  illustrate  the  practical  applications  of  generative  models  for  robot  learning.  Our  non-autoregressive,  action-conditioned  video  generation  model  can  act  as  a  world  model,  enabling  embodied  agents  to  plan  using  visual  model-predictive  control.  Furthermore,  I'll  showcase  a  generalist  agent  trained  via  next  token  prediction  to  learn  from  diverse  robotic  experiences  across  various  robots  and  tasks.
■590    ▼aSchool  code:  0212.
■650  4▼aRobots
■650  4▼aSuccess
■650  4▼aFailure  analysis
■650  4▼aVideo  recordings
■650  4▼aSemantics
■650  4▼aFilm  studies
■650  4▼aLogic
■650  4▼aRobotics
■690    ▼a0900
■690    ▼a0395
■690    ▼a0771
■71020▼aStanford  University.
■7730  ▼tDissertations  Abstracts  International▼g86-04B.
■790    ▼a0212
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17163752▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF10673 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.