본문

서브메뉴

Learning to Interact With the 3D World
Learning to Interact With the 3D World
Learning to Interact With the 3D World

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211153016
ISBN  
9798384045878
DDC  
004
저자명  
Qian, Shengyi.
서명/저자  
Learning to Interact With the 3D World
발행사항  
[Sl] : University of Michigan, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
175 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-04, Section: A.
주기사항  
Advisor: Chai, Joyce;Fouhey, David Ford.
학위논문주기  
Thesis (Ph.D.)--University of Michigan, 2024.
초록/해제  
요약Enabling machines to perceive, understand, and interact with the 3D world is a fundamental challenge in Computer Vision and Robotics, crucial for embodied agents operating in real-world environments. This dissertation addresses this challenge by proposing novel approaches to unify 3D reconstruction, affordance learning, and object manipulation. By developing a system that can understand and interact with arbitrary objects in diverse scenes, this research aims to enhance the ability of AI agents to navigate and operate in both physical and digital worlds.This dissertation is structured into three primary parts, each tackling a key challenge in 3D perception and interaction. First, we introduce novel approaches for 3D scene understanding from visual observations, including Associative3D and ViewSeg, which enable machines to perceive and understand the 3D world from limited input data.Next, we develop methodologies for learning and grounding affordances in 3D scenes, leveraging unstructured Internet videos and Vision Language Models to enhance the system's understanding of object interactions. We propose a novel approach for predicting 3D object interactions from a single image and investigate the distillation of comprehensive world knowledge from Vision Language Models. Finally, we extend our approach to active object manipulation, employing the developed system as a visual pretraining mechanism for robotics to improve the performance and generalization of manipulation policies.The key contributions of this dissertation lie in the development of techniques spanning from passive 3D perception to active object manipulation. By leveraging the vast knowledge from demonstration videos and Vision Language Models and applying the system to robotic pretraining, this research takes significant steps towards endowing machines with the ability to intelligently interact with the 3D world. The proposed methodologies collectively advance the state-of-the-art in machine perception and interaction, paving the way for more capable and adaptable AI agents.
일반주제명  
Computer science
일반주제명  
Robotics
일반주제명  
Language
일반주제명  
Information technology
키워드  
3D perception
키워드  
Computer Vision
키워드  
Affordance learning
키워드  
Embodied AI
키워드  
Real-world environments
기타저자  
University of Michigan Computer Science & Engineering
기본자료저록  
Dissertations Abstracts International. 86-04A.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017164555
■00520250211153016
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798384045878
■035    ▼a(MiAaPQ)AAI31631530
■035    ▼a(MiAaPQ)umichrackham005831
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aQian,  Shengyi.
■24510▼aLearning  to  Interact  With  the  3D  World
■260    ▼a[Sl]▼bUniversity  of  Michigan▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a175  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-04,  Section:  A.
■500    ▼aAdvisor:  Chai,  Joyce;Fouhey,  David  Ford.
■5021  ▼aThesis  (Ph.D.)--University  of  Michigan,  2024.
■520    ▼aEnabling  machines  to  perceive,  understand,  and  interact  with  the  3D  world  is  a  fundamental  challenge  in  Computer  Vision  and  Robotics,  crucial  for  embodied  agents  operating  in  real-world  environments.  This  dissertation  addresses  this  challenge  by  proposing  novel  approaches  to  unify  3D  reconstruction,  affordance  learning,  and  object  manipulation.  By  developing  a  system  that  can  understand  and  interact  with  arbitrary  objects  in  diverse  scenes,  this  research  aims  to  enhance  the  ability  of  AI  agents  to  navigate  and  operate  in  both  physical  and  digital  worlds.This  dissertation  is  structured  into  three  primary  parts,  each  tackling  a  key  challenge  in  3D  perception  and  interaction.  First,  we  introduce  novel  approaches  for  3D  scene  understanding  from  visual  observations,  including  Associative3D  and  ViewSeg,  which  enable  machines  to  perceive  and  understand  the  3D  world  from  limited  input  data.Next,  we  develop  methodologies  for  learning  and  grounding  affordances  in  3D  scenes,  leveraging  unstructured  Internet  videos  and  Vision  Language  Models  to  enhance  the  system's  understanding  of  object  interactions.  We  propose  a  novel  approach  for  predicting  3D  object  interactions  from  a  single  image  and  investigate  the  distillation  of  comprehensive  world  knowledge  from  Vision  Language  Models. Finally,  we  extend  our  approach  to  active  object  manipulation,  employing  the  developed  system  as  a  visual  pretraining  mechanism  for  robotics  to  improve  the  performance  and  generalization  of  manipulation  policies.The  key  contributions  of  this  dissertation  lie  in  the  development  of  techniques  spanning  from  passive  3D  perception  to  active  object  manipulation.  By  leveraging  the  vast  knowledge  from  demonstration  videos  and  Vision  Language  Models  and  applying  the  system  to  robotic  pretraining,  this  research  takes  significant  steps  towards  endowing  machines  with  the  ability  to  intelligently  interact  with  the  3D  world.  The  proposed  methodologies  collectively  advance  the  state-of-the-art  in  machine  perception  and  interaction,  paving  the  way  for  more  capable  and  adaptable  AI  agents.
■590    ▼aSchool  code:  0127.
■650  4▼aComputer  science
■650  4▼aRobotics
■650  4▼aLanguage
■650  4▼aInformation  technology
■653    ▼a3D  perception
■653    ▼aComputer  Vision
■653    ▼aAffordance  learning
■653    ▼aEmbodied  AI
■653    ▼aReal-world  environments
■690    ▼a0984
■690    ▼a0800
■690    ▼a0771
■690    ▼a0489
■690    ▼a0679
■71020▼aUniversity  of  Michigan▼bComputer  Science  &  Engineering.
■7730  ▼tDissertations  Abstracts  International▼g86-04A.
■790    ▼a0127
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164555▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF13591 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.