서브메뉴
검색
Spatial Reasoning in Dynamic Scenes
Spatial Reasoning in Dynamic Scenes
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211153026
- ISBN
- 9798342744997
- DDC
- 621.3
- 서명/저자
- Spatial Reasoning in Dynamic Scenes
- 발행사항
- [Sl] : Columbia University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 127 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-05, Section: B.
- 주기사항
- Advisor: Vondrick, Carl.
- 학위논문주기
- Thesis (Ph.D.)--Columbia University, 2024.
- 초록/해제
- 요약Over the past several years, machine learning has enabled incredible progress on many tasks, such as mastering board games, recognizing objects, conversing in natural language, and generating images or videos. Despite these accomplishments, state-of-the-art techniques in artificial intelligence lack the foundations necessary to flexibly and robustly understand and manipulate their three-dimensional spatial surroundings. For instance, before their second birthday, children learn that objects persist during occlusion, they know how containment works, and they are surprised by novel physics. In contrast, a true notion of object permanence has remained elusive for computer vision, despite its vitality in perceiving and interacting with everyday situations. In this thesis, I will outline my work on enhancing spatial reasoning within dynamic scenes, where I have integrated machine learning, intuitive physics, geometry, and world knowledge to create powerful frameworks that can capture, represent, and generate their complex, cluttered visual environment. Specifically, I will present models to reconstruct 4D scenes, track objects through occlusions, and perform dynamic view synthesis, all from a single camera viewpoint, and often successfully generalizing to real-world settings. These capabilities are pivotal for applications in embodied intelligence (such as robotics and self-driving), content creation and editing, or augmented and mixed reality, where machines need to accurately represent their surroundings and deeply understand how they evolve over time.
- 일반주제명
- Computer engineering
- 키워드
- Computer vision
- 키워드
- Deep learning
- 키워드
- Object tracking
- 기타저자
- Columbia University Computer Science
- 기본자료저록
- Dissertations Abstracts International. 86-05B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017164638
■00520250211153026
■006m o d
■007cr#unu||||||||
■020 ▼a9798342744997
■035 ▼a(MiAaPQ)AAI31633912
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a621.3
■1001 ▼aVan Hoorick, Basile.
■24510▼aSpatial Reasoning in Dynamic Scenes
■260 ▼a[Sl]▼bColumbia University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a127 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-05, Section: B.
■500 ▼aAdvisor: Vondrick, Carl.
■5021 ▼aThesis (Ph.D.)--Columbia University, 2024.
■520 ▼aOver the past several years, machine learning has enabled incredible progress on many tasks, such as mastering board games, recognizing objects, conversing in natural language, and generating images or videos. Despite these accomplishments, state-of-the-art techniques in artificial intelligence lack the foundations necessary to flexibly and robustly understand and manipulate their three-dimensional spatial surroundings. For instance, before their second birthday, children learn that objects persist during occlusion, they know how containment works, and they are surprised by novel physics. In contrast, a true notion of object permanence has remained elusive for computer vision, despite its vitality in perceiving and interacting with everyday situations. In this thesis, I will outline my work on enhancing spatial reasoning within dynamic scenes, where I have integrated machine learning, intuitive physics, geometry, and world knowledge to create powerful frameworks that can capture, represent, and generate their complex, cluttered visual environment. Specifically, I will present models to reconstruct 4D scenes, track objects through occlusions, and perform dynamic view synthesis, all from a single camera viewpoint, and often successfully generalizing to real-world settings. These capabilities are pivotal for applications in embodied intelligence (such as robotics and self-driving), content creation and editing, or augmented and mixed reality, where machines need to accurately represent their surroundings and deeply understand how they evolve over time.
■590 ▼aSchool code: 0054.
■650 4▼aComputer engineering
■653 ▼a3D reconstruction
■653 ▼aComputer vision
■653 ▼aDeep learning
■653 ▼aObject tracking
■653 ▼aScene understanding
■653 ▼aSpatial reasoning
■690 ▼a0800
■690 ▼a0464
■71020▼aColumbia University▼bComputer Science.
■7730 ▼tDissertations Abstracts International▼g86-05B.
■790 ▼a0054
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164638▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


