본문

서브메뉴

Bridging Understanding and Generation: A Vision Perspective
Bridging Understanding and Generation: A Vision Perspective
Bridging Understanding and Generation: A Vision Perspective

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202102958
ISBN  
9798291590423
DDC  
004
저자명  
Li, Xiang.
서명/저자  
Bridging Understanding and Generation: A Vision Perspective
발행사항  
[Sl] : Carnegie Mellon University, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
104 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-03, Section: B.
주기사항  
Advisor: Ramakrishnan, Bhiksha.
학위논문주기  
Thesis (Ph.D.)--Carnegie Mellon University, 2025.
초록/해제  
요약Understanding and generation are cornerstone capabilities of visual intelligence, enabling systems to interpret complex visual scenes and construct meaningful representations. Advanced understanding models exhibit remarkable proficiency in comprehending scenes, even under challenging conditions. Concurrently, generative models, such as diffusion models and autoregressive models, have demonstrated impressive zero-shot capabilities, generating photorealistic images for diverse and intricate scenarios.Despite these advancements, the interplay between visual understanding and generation remains underexplored. Visual understanding extracts high-level semantics from raw RGB images, creating compact and meaningful representations of visual scenes. Conversely, visual generation decodes these compact representations back into realistic RGB images. Bridging these two domains presents an opportunity to foster a mutually beneficial relationship, leveraging their inherent complementarities.This thesis seeks to bridge the gap between visual understanding and generation by exploring how these domains can complement and enhance one another. The work begins by analyzing and validating the individual effectiveness of understanding and generation models. It then focuses on integrating these domains, revealing the underlying relationships and synergies between them. The key contributions of this thesis are as follows:We present studies demonstrating how visual generation can benefit from visual understanding and vice versa, leveraging shared knowledge from learned representations. We explore integrating understanding and generation within a unified generative framework, enhancing performance by enriching the model's latent space. We conduct comprehensive experiments in heterogeneous settings to evaluate the impact of architectural design choices, modalities, and training methodologies. This thesis provides valuable insights into intelligent multimedia analysis in the era of deep learning, with practical implications for multimodal forensic understanding and deduction. It aspires to inspire further research in related fields, advancing the frontiers of visual intelligence.
일반주제명  
Computer science
일반주제명  
Computer engineering
일반주제명  
Electrical engineering
키워드  
Cornerstone capabilities
키워드  
Visual intelligence
키워드  
Zero-shot capabilities
기타저자  
Carnegie Mellon University Electrical and Computer Engineering
기본자료저록  
Dissertations Abstracts International. 87-03B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017356587
■00520260202102958
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798291590423
■035    ▼a(MiAaPQ)AAI31770661
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aLi,  Xiang.
■24510▼aBridging  Understanding  and  Generation:  A  Vision  Perspective
■260    ▼a[Sl]▼bCarnegie  Mellon  University▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a104  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-03,  Section:  B.
■500    ▼aAdvisor:  Ramakrishnan,  Bhiksha.
■5021  ▼aThesis  (Ph.D.)--Carnegie  Mellon  University,  2025.
■520    ▼aUnderstanding  and  generation  are  cornerstone  capabilities  of  visual  intelligence,  enabling  systems  to  interpret  complex  visual  scenes  and  construct  meaningful  representations.  Advanced  understanding  models  exhibit  remarkable  proficiency  in  comprehending  scenes,  even  under  challenging  conditions.  Concurrently,  generative  models,  such  as  diffusion  models  and  autoregressive  models,  have  demonstrated  impressive  zero-shot  capabilities,  generating  photorealistic  images  for  diverse  and  intricate  scenarios.Despite  these  advancements,  the  interplay  between  visual  understanding  and  generation  remains  underexplored.  Visual  understanding  extracts  high-level  semantics  from  raw  RGB  images,  creating  compact  and  meaningful  representations  of  visual  scenes.  Conversely,  visual  generation  decodes  these  compact  representations  back  into  realistic  RGB  images.  Bridging  these  two  domains  presents  an  opportunity  to  foster  a  mutually  beneficial  relationship,  leveraging  their  inherent  complementarities.This  thesis  seeks  to  bridge  the  gap  between  visual  understanding  and  generation  by  exploring  how  these  domains  can  complement  and  enhance  one  another.  The  work  begins  by  analyzing  and  validating  the  individual  effectiveness  of  understanding  and  generation  models.  It  then  focuses  on  integrating  these  domains,  revealing  the  underlying  relationships  and  synergies  between  them.  The  key  contributions  of  this  thesis  are  as  follows:We  present  studies  demonstrating  how  visual  generation  can  benefit  from  visual  understanding  and  vice  versa,  leveraging  shared  knowledge  from  learned  representations.  We  explore  integrating  understanding  and  generation  within  a  unified  generative  framework,  enhancing  performance  by  enriching  the  model's  latent  space.  We  conduct  comprehensive  experiments  in  heterogeneous  settings  to  evaluate  the  impact  of  architectural  design  choices,  modalities,  and  training  methodologies.  This  thesis  provides  valuable  insights  into  intelligent  multimedia  analysis  in  the  era  of  deep  learning,  with  practical  implications  for  multimodal  forensic  understanding  and  deduction.  It  aspires  to  inspire  further  research  in  related  fields,  advancing  the  frontiers  of  visual  intelligence.
■590    ▼aSchool  code:  0041.
■650  4▼aComputer  science
■650  4▼aComputer  engineering
■650  4▼aElectrical  engineering
■653    ▼aCornerstone  capabilities
■653    ▼aVisual  intelligence
■653    ▼aZero-shot  capabilities
■690    ▼a0984
■690    ▼a0544
■690    ▼a0464
■71020▼aCarnegie  Mellon  University▼bElectrical  and  Computer  Engineering.
■7730  ▼tDissertations  Abstracts  International▼g87-03B.
■790    ▼a0041
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17356587▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF15338 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.