서브메뉴
검색
Bridging Understanding and Generation: A Vision Perspective
Bridging Understanding and Generation: A Vision Perspective
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202102958
- ISBN
- 9798291590423
- DDC
- 004
- 저자명
- Li, Xiang.
- 서명/저자
- Bridging Understanding and Generation: A Vision Perspective
- 발행사항
- [Sl] : Carnegie Mellon University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 104 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-03, Section: B.
- 주기사항
- Advisor: Ramakrishnan, Bhiksha.
- 학위논문주기
- Thesis (Ph.D.)--Carnegie Mellon University, 2025.
- 초록/해제
- 요약Understanding and generation are cornerstone capabilities of visual intelligence, enabling systems to interpret complex visual scenes and construct meaningful representations. Advanced understanding models exhibit remarkable proficiency in comprehending scenes, even under challenging conditions. Concurrently, generative models, such as diffusion models and autoregressive models, have demonstrated impressive zero-shot capabilities, generating photorealistic images for diverse and intricate scenarios.Despite these advancements, the interplay between visual understanding and generation remains underexplored. Visual understanding extracts high-level semantics from raw RGB images, creating compact and meaningful representations of visual scenes. Conversely, visual generation decodes these compact representations back into realistic RGB images. Bridging these two domains presents an opportunity to foster a mutually beneficial relationship, leveraging their inherent complementarities.This thesis seeks to bridge the gap between visual understanding and generation by exploring how these domains can complement and enhance one another. The work begins by analyzing and validating the individual effectiveness of understanding and generation models. It then focuses on integrating these domains, revealing the underlying relationships and synergies between them. The key contributions of this thesis are as follows:We present studies demonstrating how visual generation can benefit from visual understanding and vice versa, leveraging shared knowledge from learned representations. We explore integrating understanding and generation within a unified generative framework, enhancing performance by enriching the model's latent space. We conduct comprehensive experiments in heterogeneous settings to evaluate the impact of architectural design choices, modalities, and training methodologies. This thesis provides valuable insights into intelligent multimedia analysis in the era of deep learning, with practical implications for multimodal forensic understanding and deduction. It aspires to inspire further research in related fields, advancing the frontiers of visual intelligence.
- 일반주제명
- Computer science
- 일반주제명
- Computer engineering
- 일반주제명
- Electrical engineering
- 기타저자
- Carnegie Mellon University Electrical and Computer Engineering
- 기본자료저록
- Dissertations Abstracts International. 87-03B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017356587
■00520260202102958
■006m o d
■007cr#unu||||||||
■020 ▼a9798291590423
■035 ▼a(MiAaPQ)AAI31770661
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aLi, Xiang.
■24510▼aBridging Understanding and Generation: A Vision Perspective
■260 ▼a[Sl]▼bCarnegie Mellon University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a104 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-03, Section: B.
■500 ▼aAdvisor: Ramakrishnan, Bhiksha.
■5021 ▼aThesis (Ph.D.)--Carnegie Mellon University, 2025.
■520 ▼aUnderstanding and generation are cornerstone capabilities of visual intelligence, enabling systems to interpret complex visual scenes and construct meaningful representations. Advanced understanding models exhibit remarkable proficiency in comprehending scenes, even under challenging conditions. Concurrently, generative models, such as diffusion models and autoregressive models, have demonstrated impressive zero-shot capabilities, generating photorealistic images for diverse and intricate scenarios.Despite these advancements, the interplay between visual understanding and generation remains underexplored. Visual understanding extracts high-level semantics from raw RGB images, creating compact and meaningful representations of visual scenes. Conversely, visual generation decodes these compact representations back into realistic RGB images. Bridging these two domains presents an opportunity to foster a mutually beneficial relationship, leveraging their inherent complementarities.This thesis seeks to bridge the gap between visual understanding and generation by exploring how these domains can complement and enhance one another. The work begins by analyzing and validating the individual effectiveness of understanding and generation models. It then focuses on integrating these domains, revealing the underlying relationships and synergies between them. The key contributions of this thesis are as follows:We present studies demonstrating how visual generation can benefit from visual understanding and vice versa, leveraging shared knowledge from learned representations. We explore integrating understanding and generation within a unified generative framework, enhancing performance by enriching the model's latent space. We conduct comprehensive experiments in heterogeneous settings to evaluate the impact of architectural design choices, modalities, and training methodologies. This thesis provides valuable insights into intelligent multimedia analysis in the era of deep learning, with practical implications for multimodal forensic understanding and deduction. It aspires to inspire further research in related fields, advancing the frontiers of visual intelligence.
■590 ▼aSchool code: 0041.
■650 4▼aComputer science
■650 4▼aComputer engineering
■650 4▼aElectrical engineering
■653 ▼aCornerstone capabilities
■653 ▼aVisual intelligence
■653 ▼aZero-shot capabilities
■690 ▼a0984
■690 ▼a0544
■690 ▼a0464
■71020▼aCarnegie Mellon University▼bElectrical and Computer Engineering.
■7730 ▼tDissertations Abstracts International▼g87-03B.
■790 ▼a0041
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17356587▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


