서브메뉴
검색
Learning to Interact With the 3D World
Learning to Interact With the 3D World
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211153016
- ISBN
- 9798384045878
- DDC
- 004
- 저자명
- Qian, Shengyi.
- 서명/저자
- Learning to Interact With the 3D World
- 발행사항
- [Sl] : University of Michigan, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 175 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-04, Section: A.
- 주기사항
- Advisor: Chai, Joyce;Fouhey, David Ford.
- 학위논문주기
- Thesis (Ph.D.)--University of Michigan, 2024.
- 초록/해제
- 요약Enabling machines to perceive, understand, and interact with the 3D world is a fundamental challenge in Computer Vision and Robotics, crucial for embodied agents operating in real-world environments. This dissertation addresses this challenge by proposing novel approaches to unify 3D reconstruction, affordance learning, and object manipulation. By developing a system that can understand and interact with arbitrary objects in diverse scenes, this research aims to enhance the ability of AI agents to navigate and operate in both physical and digital worlds.This dissertation is structured into three primary parts, each tackling a key challenge in 3D perception and interaction. First, we introduce novel approaches for 3D scene understanding from visual observations, including Associative3D and ViewSeg, which enable machines to perceive and understand the 3D world from limited input data.Next, we develop methodologies for learning and grounding affordances in 3D scenes, leveraging unstructured Internet videos and Vision Language Models to enhance the system's understanding of object interactions. We propose a novel approach for predicting 3D object interactions from a single image and investigate the distillation of comprehensive world knowledge from Vision Language Models. Finally, we extend our approach to active object manipulation, employing the developed system as a visual pretraining mechanism for robotics to improve the performance and generalization of manipulation policies.The key contributions of this dissertation lie in the development of techniques spanning from passive 3D perception to active object manipulation. By leveraging the vast knowledge from demonstration videos and Vision Language Models and applying the system to robotic pretraining, this research takes significant steps towards endowing machines with the ability to intelligently interact with the 3D world. The proposed methodologies collectively advance the state-of-the-art in machine perception and interaction, paving the way for more capable and adaptable AI agents.
- 일반주제명
- Computer science
- 일반주제명
- Robotics
- 일반주제명
- Language
- 일반주제명
- Information technology
- 키워드
- 3D perception
- 키워드
- Computer Vision
- 키워드
- Embodied AI
- 기타저자
- University of Michigan Computer Science & Engineering
- 기본자료저록
- Dissertations Abstracts International. 86-04A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017164555
■00520250211153016
■006m o d
■007cr#unu||||||||
■020 ▼a9798384045878
■035 ▼a(MiAaPQ)AAI31631530
■035 ▼a(MiAaPQ)umichrackham005831
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aQian, Shengyi.
■24510▼aLearning to Interact With the 3D World
■260 ▼a[Sl]▼bUniversity of Michigan▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a175 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-04, Section: A.
■500 ▼aAdvisor: Chai, Joyce;Fouhey, David Ford.
■5021 ▼aThesis (Ph.D.)--University of Michigan, 2024.
■520 ▼aEnabling machines to perceive, understand, and interact with the 3D world is a fundamental challenge in Computer Vision and Robotics, crucial for embodied agents operating in real-world environments. This dissertation addresses this challenge by proposing novel approaches to unify 3D reconstruction, affordance learning, and object manipulation. By developing a system that can understand and interact with arbitrary objects in diverse scenes, this research aims to enhance the ability of AI agents to navigate and operate in both physical and digital worlds.This dissertation is structured into three primary parts, each tackling a key challenge in 3D perception and interaction. First, we introduce novel approaches for 3D scene understanding from visual observations, including Associative3D and ViewSeg, which enable machines to perceive and understand the 3D world from limited input data.Next, we develop methodologies for learning and grounding affordances in 3D scenes, leveraging unstructured Internet videos and Vision Language Models to enhance the system's understanding of object interactions. We propose a novel approach for predicting 3D object interactions from a single image and investigate the distillation of comprehensive world knowledge from Vision Language Models. Finally, we extend our approach to active object manipulation, employing the developed system as a visual pretraining mechanism for robotics to improve the performance and generalization of manipulation policies.The key contributions of this dissertation lie in the development of techniques spanning from passive 3D perception to active object manipulation. By leveraging the vast knowledge from demonstration videos and Vision Language Models and applying the system to robotic pretraining, this research takes significant steps towards endowing machines with the ability to intelligently interact with the 3D world. The proposed methodologies collectively advance the state-of-the-art in machine perception and interaction, paving the way for more capable and adaptable AI agents.
■590 ▼aSchool code: 0127.
■650 4▼aComputer science
■650 4▼aRobotics
■650 4▼aLanguage
■650 4▼aInformation technology
■653 ▼a3D perception
■653 ▼aComputer Vision
■653 ▼aAffordance learning
■653 ▼aEmbodied AI
■653 ▼aReal-world environments
■690 ▼a0984
■690 ▼a0800
■690 ▼a0771
■690 ▼a0489
■690 ▼a0679
■71020▼aUniversity of Michigan▼bComputer Science & Engineering.
■7730 ▼tDissertations Abstracts International▼g86-04A.
■790 ▼a0127
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164555▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


