서브메뉴
검색
Discovering the 4D World Behind Any Video
Discovering the 4D World Behind Any Video
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211152736
- ISBN
- 9798384448730
- DDC
- 004
- 저자명
- Ye, Vickie.
- 서명/저자
- Discovering the 4D World Behind Any Video
- 발행사항
- [Sl] : University of California, Berkeley, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 118 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-03, Section: B.
- 주기사항
- Advisor: Kanazawa, Angjoo.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Berkeley, 2024.
- 초록/해제
- 요약As we begin to interact with AI systems, we need them to be able to interpret the visual world in 4D - that is, to perceive the geometry and motion in the world. However, pixel differences in image space result from either geometry (via camera motion) or scene motion in the world. To disentangle these two sources this from a single video is extremely under-constrained.In this thesis, I build several systems that recover scene representations from limited image observations. Specifically, I study a series of problems that build toward the 4D monocular recovery problem, each one addressing a different aspect of the under-constrained nature of the problem. First I study the problem of recovering shape from under-constrained inputs, without scene motion. Specifically, I present pixelNeRF, a method to synthesize novel views of a static scene from single or few views. We learn a scene prior by training a 3D neural representation conditioned on image features across multiple scenes. This learned scene prior enables 3D scene completion from the under-constrained inputs of single or few images. Next I study the problem of recovering motion without 3D shape. In particular, I present Deformable Sprites, a method to extract persistent elements of a dynamic scene from an input video. We represent each element as 2D image layers that deform across the video.Finally I present two studies of performing the joint recovery of both the shape and motion of the 4D world from any single video. I first study the special case of dynamic humans, and present SLAHMR, in which we recover from a single video the global poses of all the humans and the camera in the world coordinate frame. I then move on to the general case of recovering any dynamic objects from a single video in Shape of Motion, in which we recover the entire scene as 4D gaussians, which we can use for dynamic novel view synthesis and 3D tracking.
- 일반주제명
- Computer science
- 일반주제명
- Computer engineering
- 일반주제명
- Information technology
- 키워드
- Computer vision
- 기타저자
- University of California, Berkeley Electrical Engineering & Computer Sciences
- 기본자료저록
- Dissertations Abstracts International. 86-03B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017163649
■00520250211152736
■006m o d
■007cr#unu||||||||
■020 ▼a9798384448730
■035 ▼a(MiAaPQ)AAI31491367
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aYe, Vickie.
■24510▼aDiscovering the 4D World Behind Any Video
■260 ▼a[Sl]▼bUniversity of California, Berkeley▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a118 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-03, Section: B.
■500 ▼aAdvisor: Kanazawa, Angjoo.
■5021 ▼aThesis (Ph.D.)--University of California, Berkeley, 2024.
■520 ▼aAs we begin to interact with AI systems, we need them to be able to interpret the visual world in 4D - that is, to perceive the geometry and motion in the world. However, pixel differences in image space result from either geometry (via camera motion) or scene motion in the world. To disentangle these two sources this from a single video is extremely under-constrained.In this thesis, I build several systems that recover scene representations from limited image observations. Specifically, I study a series of problems that build toward the 4D monocular recovery problem, each one addressing a different aspect of the under-constrained nature of the problem. First I study the problem of recovering shape from under-constrained inputs, without scene motion. Specifically, I present pixelNeRF, a method to synthesize novel views of a static scene from single or few views. We learn a scene prior by training a 3D neural representation conditioned on image features across multiple scenes. This learned scene prior enables 3D scene completion from the under-constrained inputs of single or few images. Next I study the problem of recovering motion without 3D shape. In particular, I present Deformable Sprites, a method to extract persistent elements of a dynamic scene from an input video. We represent each element as 2D image layers that deform across the video.Finally I present two studies of performing the joint recovery of both the shape and motion of the 4D world from any single video. I first study the special case of dynamic humans, and present SLAHMR, in which we recover from a single video the global poses of all the humans and the camera in the world coordinate frame. I then move on to the general case of recovering any dynamic objects from a single video in Shape of Motion, in which we recover the entire scene as 4D gaussians, which we can use for dynamic novel view synthesis and 3D tracking.
■590 ▼aSchool code: 0028.
■650 4▼aComputer science
■650 4▼aComputer engineering
■650 4▼aInformation technology
■653 ▼a3D reconstruction
■653 ▼aComputer graphics
■653 ▼aComputer vision
■653 ▼aVideo understanding
■653 ▼a4D monocular recovery problem
■690 ▼a0984
■690 ▼a0800
■690 ▼a0489
■690 ▼a0464
■71020▼aUniversity of California, Berkeley▼bElectrical Engineering & Computer Sciences.
■7730 ▼tDissertations Abstracts International▼g86-03B.
■790 ▼a0028
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17163649▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


