서브메뉴
검색
From 3D Mapping to Scene Representations for Embodied AI
From 3D Mapping to Scene Representations for Embodied AI
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260209102904
- ISBN
- 9798265405449
- DDC
- 616.8
- 서명/저자
- From 3D Mapping to Scene Representations for Embodied AI
- 발행사항
- [Sl] : Georgia Institute of Technology, 2023
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2023
- 형태사항
- 153 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-06, Section: B.
- 주기사항
- Advisor: Essa, Irfan;Romberg, Justin.
- 학위논문주기
- Thesis (Ph.D.)--Georgia Institute of Technology, 2023.
- 초록/해제
- 요약In the past few years, a burgeoning field of research has emerged within the broader AI community known as "Embodied AI." This field encompasses various challenges, including the development of scene datasets and simulators used to train AI agents in diverse tasks, necessitating a comprehensive set of skills. Generally speaking, Embodied AI research projects assume a similar setup where an agent equipped with a set of sensors (usually an RGB and Depth camera) is trained to accomplish certain tasks such as navigation, pick-and-place, question answering etc. While there are many approaches and designs of such Embodied AI agents, this thesis focuses on methods involving intermediate scene representations. In contrast to end-to-end approaches, AI systems with intermediate explicit scene representations comprise two distinct modules: one for raw input sensor processing and a second one for planning and acting. Agents utilizing scene representations offer several advantages, including reduced susceptibility to the forgetting effect observed in their RNN counterparts, easier incorporation of inductive bias directly into the representations (e.g., geometrical constraints), and improved interpretability.Towards this end, this thesis serves as an investigation towards the design of such scene representations. We specifically research how to leverage 3D mapping techniques in order to build rich, dense and useful representations for Embodied AI applications. We start by studying representations in the form of 2D topdown metric maps. These 2D maps can store features, labels or geometrical information to form a useful training signal for the tasks of semantic mapping, navigation or question answering. We then study 3D representations in the forms of implicit maps applied for SLAM and 3D object-based maps for multi-object re-identification. Next we extend our research to dynamic scenes and explore 4D representations in the form of localized keyframes. Finally, the thesis also explores connections between these different representations while highlighting their strengths and limitations.
- 일반주제명
- Episodic memory
- 일반주제명
- Semantics
- 기본자료저록
- Dissertations Abstracts International. 87-06B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260203s2023 us c eng d■001000017365961
■00520260209102904
■006m o d
■007cr#unu||||||||
■020 ▼a9798265405449
■035 ▼a(MiAaPQ)AAI32315571
■035 ▼a(MiAaPQ)GeorgiaTech73112
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a616.8
■1001 ▼aCartillier, Vincent.
■24510▼aFrom 3D Mapping to Scene Representations for Embodied AI
■260 ▼a[Sl]▼bGeorgia Institute of Technology▼c2023
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2023
■300 ▼a153 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-06, Section: B.
■500 ▼aAdvisor: Essa, Irfan;Romberg, Justin.
■5021 ▼aThesis (Ph.D.)--Georgia Institute of Technology, 2023.
■520 ▼aIn the past few years, a burgeoning field of research has emerged within the broader AI community known as "Embodied AI." This field encompasses various challenges, including the development of scene datasets and simulators used to train AI agents in diverse tasks, necessitating a comprehensive set of skills. Generally speaking, Embodied AI research projects assume a similar setup where an agent equipped with a set of sensors (usually an RGB and Depth camera) is trained to accomplish certain tasks such as navigation, pick-and-place, question answering etc. While there are many approaches and designs of such Embodied AI agents, this thesis focuses on methods involving intermediate scene representations. In contrast to end-to-end approaches, AI systems with intermediate explicit scene representations comprise two distinct modules: one for raw input sensor processing and a second one for planning and acting. Agents utilizing scene representations offer several advantages, including reduced susceptibility to the forgetting effect observed in their RNN counterparts, easier incorporation of inductive bias directly into the representations (e.g., geometrical constraints), and improved interpretability.Towards this end, this thesis serves as an investigation towards the design of such scene representations. We specifically research how to leverage 3D mapping techniques in order to build rich, dense and useful representations for Embodied AI applications. We start by studying representations in the form of 2D topdown metric maps. These 2D maps can store features, labels or geometrical information to form a useful training signal for the tasks of semantic mapping, navigation or question answering. We then study 3D representations in the forms of implicit maps applied for SLAM and 3D object-based maps for multi-object re-identification. Next we extend our research to dynamic scenes and explore 4D representations in the form of localized keyframes. Finally, the thesis also explores connections between these different representations while highlighting their strengths and limitations.
■590 ▼aSchool code: 0078.
■650 4▼aEpisodic memory
■650 4▼aSemantics
■653 ▼a3D mapping techniques
■653 ▼aGeometrical information
■690 ▼a0800
■71020▼aGeorgia Institute of Technology.
■7730 ▼tDissertations Abstracts International▼g87-06B.
■790 ▼a0078
■791 ▼aPh.D.
■792 ▼a2023
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17365961▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


