서브메뉴
검색
Leveraging Human Videos for Robot Learning
Leveraging Human Videos for Robot Learning
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105619
- ISBN
- 9798265428486
- DDC
- 620
- 서명/저자
- Leveraging Human Videos for Robot Learning
- 발행사항
- [Sl] : Stanford University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 91 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-05, Section: A.
- 주기사항
- Advisor: Bohg, Jeannette;Cutkosky, Mark.
- 학위논문주기
- Thesis (Ph.D.)--Stanford University, 2025.
- 초록/해제
- 요약Robot learning faces a fundamental data scarcity problem that limits the development of generalpurpose manipulation policies. While other AI domains benefit from vast internet-scale datasets, robotics relies on expensive, time-consuming robot demonstrations that are dicult to scale across diverse environments and tasks. This dissertation addresses this challenge by demonstrating that human videos-abundant and diverse-can serve as an e↵ective additional training resource for robot policies when the visual embodiment gap between humans and robots is explicitly closed.We present a systematic approach across three increasingly challenging settings. First, we introduce Shadow, a method for robot-to-robot policy transfer that uses composite segmentation masks to align visual appearances across di↵erent robot embodiments, achieving over 2⇥ improvement in cross-embodiment transfer without requiring target robot data. Second, we develop Phantom, which enables zero-shot robot policy learning from hand-collected human videos by extracting actions through hand pose estimation and replacing human arms with rendered robots, achieving up to 92% success rates on diverse manipulation tasks. Finally, we present Masquerade, which scales to in-the-wild human videos by editing 675K video frames and using them to pretrain vision representations, resulting in 5-6⇥ performance improvements on long-horizon bimanual tasks compared to baseline methods.Our key insight is that explicit visual alignment-systematically reducing the distribution shift between human and robot observation spaces-enables models to e↵ectively leverage human demonstration data at scale. Across all three methods, we demonstrate that data editing techniques that bridge embodiment gaps unlock significant performance gains, suggesting a path toward leveraging the vast repository of human manipulation data available online for scalable robot learning. This work establishes explicitly bridging the visual embodiment gap as a fundamental principle for crossembodiment policy transfer and opens new pathways for training robust robot policies without the traditional constraints of robot-specific data collection.
- 일반주제명
- Robots
- 일반주제명
- Video recordings
- 일반주제명
- Robotics
- 일반주제명
- Film studies
- 기타저자
- Stanford University.
- 기본자료저록
- Dissertations Abstracts International. 87-05A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017360784
■00520260202105619
■006m o d
■007cr#unu||||||||
■020 ▼a9798265428486
■035 ▼a(MiAaPQ)AAI32316487
■035 ▼a(MiAaPQ)Stanfordkj511wb9899
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a620
■1001 ▼aLepert, Marion Marcelle.
■24510▼aLeveraging Human Videos for Robot Learning
■260 ▼a[Sl]▼bStanford University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a91 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-05, Section: A.
■500 ▼aAdvisor: Bohg, Jeannette;Cutkosky, Mark.
■5021 ▼aThesis (Ph.D.)--Stanford University, 2025.
■520 ▼aRobot learning faces a fundamental data scarcity problem that limits the development of generalpurpose manipulation policies. While other AI domains benefit from vast internet-scale datasets, robotics relies on expensive, time-consuming robot demonstrations that are dicult to scale across diverse environments and tasks. This dissertation addresses this challenge by demonstrating that human videos-abundant and diverse-can serve as an e↵ective additional training resource for robot policies when the visual embodiment gap between humans and robots is explicitly closed.We present a systematic approach across three increasingly challenging settings. First, we introduce Shadow, a method for robot-to-robot policy transfer that uses composite segmentation masks to align visual appearances across di↵erent robot embodiments, achieving over 2⇥ improvement in cross-embodiment transfer without requiring target robot data. Second, we develop Phantom, which enables zero-shot robot policy learning from hand-collected human videos by extracting actions through hand pose estimation and replacing human arms with rendered robots, achieving up to 92% success rates on diverse manipulation tasks. Finally, we present Masquerade, which scales to in-the-wild human videos by editing 675K video frames and using them to pretrain vision representations, resulting in 5-6⇥ performance improvements on long-horizon bimanual tasks compared to baseline methods.Our key insight is that explicit visual alignment-systematically reducing the distribution shift between human and robot observation spaces-enables models to e↵ectively leverage human demonstration data at scale. Across all three methods, we demonstrate that data editing techniques that bridge embodiment gaps unlock significant performance gains, suggesting a path toward leveraging the vast repository of human manipulation data available online for scalable robot learning. This work establishes explicitly bridging the visual embodiment gap as a fundamental principle for crossembodiment policy transfer and opens new pathways for training robust robot policies without the traditional constraints of robot-specific data collection.
■590 ▼aSchool code: 0212.
■650 4▼aRobots
■650 4▼aVideo recordings
■650 4▼aRobotics
■650 4▼aFilm studies
■690 ▼a0771
■690 ▼a0900
■71020▼aStanford University.
■7730 ▼tDissertations Abstracts International▼g87-05A.
■790 ▼a0212
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17360784▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


