서브메뉴
검색
Generative Models of Vision and Action
Generative Models of Vision and Action
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211152748
- ISBN
- 9798342107396
- DDC
- 620
- 저자명
- Gupta, Agrim.
- 서명/저자
- Generative Models of Vision and Action
- 발행사항
- [Sl] : Stanford University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 129 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-04, Section: B.
- 주기사항
- Advisor: Li, Fei-Fei.
- 학위논문주기
- Thesis (Ph.D.)--Stanford University, 2024.
- 초록/해제
- 요약Animals and humans display remarkable ability at building internal representations of the world and using them to simulate, evaluate and select among different possible actions. This capability is learnt primarily from observation and without any supervision. Endowing autonomous agents with similar capabilities is a fundamental challenge in machine learning. In this thesis I will explore new algorithms that enable scalable representation learning from videos via prediction, generative models of visual data and their applications to robotics.To begin, I will discuss the challenges associated with using predictive learning objectives to learn visual representations. I'll introduce a simple predictive learning architecture and objective that enables learning visual representations capable of solving a wide range of visual correspondence tasks in a zero-shot manner. Subsequently, I'll present a transformer-based approach for photorealistic video generation via diffusion modeling. Our approach jointly compresses images and videos within a unified latent space, enabling training and generation across modalities. Finally, I will illustrate the practical applications of generative models for robot learning. Our non-autoregressive, action-conditioned video generation model can act as a world model, enabling embodied agents to plan using visual model-predictive control. Furthermore, I'll showcase a generalist agent trained via next token prediction to learn from diverse robotic experiences across various robots and tasks.
- 일반주제명
- Robots
- 일반주제명
- Success
- 일반주제명
- Failure analysis
- 일반주제명
- Video recordings
- 일반주제명
- Semantics
- 일반주제명
- Film studies
- 일반주제명
- Logic
- 일반주제명
- Robotics
- 기타저자
- Stanford University.
- 기본자료저록
- Dissertations Abstracts International. 86-04B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017163752
■00520250211152748
■006m o d
■007cr#unu||||||||
■020 ▼a9798342107396
■035 ▼a(MiAaPQ)AAI31520329
■035 ▼a(MiAaPQ)Stanfordwd022wx6061
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a620
■1001 ▼aGupta, Agrim.
■24510▼aGenerative Models of Vision and Action
■260 ▼a[Sl]▼bStanford University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a129 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-04, Section: B.
■500 ▼aAdvisor: Li, Fei-Fei.
■5021 ▼aThesis (Ph.D.)--Stanford University, 2024.
■520 ▼aAnimals and humans display remarkable ability at building internal representations of the world and using them to simulate, evaluate and select among different possible actions. This capability is learnt primarily from observation and without any supervision. Endowing autonomous agents with similar capabilities is a fundamental challenge in machine learning. In this thesis I will explore new algorithms that enable scalable representation learning from videos via prediction, generative models of visual data and their applications to robotics.To begin, I will discuss the challenges associated with using predictive learning objectives to learn visual representations. I'll introduce a simple predictive learning architecture and objective that enables learning visual representations capable of solving a wide range of visual correspondence tasks in a zero-shot manner. Subsequently, I'll present a transformer-based approach for photorealistic video generation via diffusion modeling. Our approach jointly compresses images and videos within a unified latent space, enabling training and generation across modalities. Finally, I will illustrate the practical applications of generative models for robot learning. Our non-autoregressive, action-conditioned video generation model can act as a world model, enabling embodied agents to plan using visual model-predictive control. Furthermore, I'll showcase a generalist agent trained via next token prediction to learn from diverse robotic experiences across various robots and tasks.
■590 ▼aSchool code: 0212.
■650 4▼aRobots
■650 4▼aSuccess
■650 4▼aFailure analysis
■650 4▼aVideo recordings
■650 4▼aSemantics
■650 4▼aFilm studies
■650 4▼aLogic
■650 4▼aRobotics
■690 ▼a0900
■690 ▼a0395
■690 ▼a0771
■71020▼aStanford University.
■7730 ▼tDissertations Abstracts International▼g86-04B.
■790 ▼a0212
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17163752▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


