서브메뉴
검색
Towards a Comprehensive Benchmark for Embodied AI and Robotics
Towards a Comprehensive Benchmark for Embodied AI and Robotics
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202104856
- ISBN
- 9798288814785
- DDC
- 530
- 저자명
- Li, Chengshu.
- 서명/저자
- Towards a Comprehensive Benchmark for Embodied AI and Robotics
- 발행사항
- [Sl] : Stanford University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 314 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-02, Section: B.
- 주기사항
- Advisor: Li, Fei-Fei.
- 학위논문주기
- Thesis (Ph.D.)--Stanford University, 2025.
- 초록/해제
- 요약This dissertation explores two key research directions toward enabling general-purpose embodied agents: the development of realistic, large-scale benchmarks and environments, and the design of learning frameworks-particularly action space representations-that support efficient policy learning for long-horizon mobile manipulation tasks. The first line of work establishes a closed-loop ecosystem for benchmarking and training embodied agents. Beginning with iGibson 1.0 and 2.0, we develop physically interactive 3D simulation platforms capable of supporting complex object interactions in realistic household environments. Building on these foundations, we introduce the BEHAVIOR and BEHAVIOR-1K benchmarks-comprising 100 and 1,000 everyday household activities, respectively-grounded in human time-use data, defined using a flexible logic-based language, and supported by human VR demonstrations. To enable scalable data-driven policy training, we propose MoMaGen, a demonstration generation method that synthesizes thousands of diverse trajectories from a single human demonstration. The second line of work investigates action space design as a source of inductive bias for solving long-horizon robotic tasks. We first present HRL4IN, a hierarchical reinforcement learning approach that decomposes interactive navigation via high-level end-effector goals. We then introduce ReLMoGen, a hybrid method that combines high-level exploration in spatial goal spaces with low-level motion generation for efficient execution. Finally, Chain of Code leverages large language models (LLMs) to generate executable code and pseudocode, enabling agents to blend algorithmic reasoning and commonsense inference for task completion. Together, these contributions advance the goal of building physically capable, semantically grounded, and human-aligned embodied agents.
- 일반주제명
- Physics
- 일반주제명
- Virtual reality
- 일반주제명
- Benchmarks
- 일반주제명
- Robotics
- 키워드
- Robotic tasks
- 기타저자
- Stanford University.
- 기본자료저록
- Dissertations Abstracts International. 87-02B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017359257
■00520260202104856
■006m o d
■007cr#unu||||||||
■020 ▼a9798288814785
■035 ▼a(MiAaPQ)AAI32201019
■035 ▼a(MiAaPQ)Stanfordwz861bd4280
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a530
■1001 ▼aLi, Chengshu.
■24510▼aTowards a Comprehensive Benchmark for Embodied AI and Robotics
■260 ▼a[Sl]▼bStanford University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a314 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-02, Section: B.
■500 ▼aAdvisor: Li, Fei-Fei.
■5021 ▼aThesis (Ph.D.)--Stanford University, 2025.
■520 ▼aThis dissertation explores two key research directions toward enabling general-purpose embodied agents: the development of realistic, large-scale benchmarks and environments, and the design of learning frameworks-particularly action space representations-that support efficient policy learning for long-horizon mobile manipulation tasks. The first line of work establishes a closed-loop ecosystem for benchmarking and training embodied agents. Beginning with iGibson 1.0 and 2.0, we develop physically interactive 3D simulation platforms capable of supporting complex object interactions in realistic household environments. Building on these foundations, we introduce the BEHAVIOR and BEHAVIOR-1K benchmarks-comprising 100 and 1,000 everyday household activities, respectively-grounded in human time-use data, defined using a flexible logic-based language, and supported by human VR demonstrations. To enable scalable data-driven policy training, we propose MoMaGen, a demonstration generation method that synthesizes thousands of diverse trajectories from a single human demonstration. The second line of work investigates action space design as a source of inductive bias for solving long-horizon robotic tasks. We first present HRL4IN, a hierarchical reinforcement learning approach that decomposes interactive navigation via high-level end-effector goals. We then introduce ReLMoGen, a hybrid method that combines high-level exploration in spatial goal spaces with low-level motion generation for efficient execution. Finally, Chain of Code leverages large language models (LLMs) to generate executable code and pseudocode, enabling agents to blend algorithmic reasoning and commonsense inference for task completion. Together, these contributions advance the goal of building physically capable, semantically grounded, and human-aligned embodied agents.
■590 ▼aSchool code: 0212.
■650 4▼aPhysics
■650 4▼aVirtual reality
■650 4▼aBenchmarks
■650 4▼aRobotics
■653 ▼aLarge language models
■653 ▼aReinforcement learning approach
■653 ▼aRobotic tasks
■690 ▼a0605
■690 ▼a0800
■690 ▼a0771
■71020▼aStanford University.
■7730 ▼tDissertations Abstracts International▼g87-02B.
■790 ▼a0212
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17359257▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


