본문

서브메뉴

Disentangled Visual Generative Models
Disentangled Visual Generative Models
Disentangled Visual Generative Models

Detailed Information

자료유형  
 학위논문 서양
최종처리일시  
20250211151428
ISBN  
9798384447214
DDC  
004
저자명  
Epstein, Dave.
서명/저자  
Disentangled Visual Generative Models
발행사항  
[Sl] : University of California, Berkeley, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
116 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-04, Section: B.
주기사항  
Advisor: Efros, Alexei A.
학위논문주기  
Thesis (Ph.D.)--University of California, Berkeley, 2024.
초록/해제  
요약Generative modeling promises an elegant solution to learning about high-dimensional data distributions such as images and videos - but how can we expose and utilize the rich structure these models discover? Rather than just drawing new samples, how can an agent actually harness p(x) as a source of knowledge about how our world works? This thesis explores scalable inductive biases that unlock a generative model's disentangled understanding of visual data, enabling much richer interaction and control as a result.First, I propose a representation of scenes as collections of feature "blobs", where a generative adversarial network (GAN) learns - without any labels - to bind each blob to a different object in the images it creates. This allows GANs to more gracefully model compositional scenes, in contrast to typical unconditional models which are constrained to highly-aligned single-object data. The trained model's representation can easily be modified to counterfactually manipulate objects in both generated and real images. Next, I consider methods that do not impose bottlenecks on architectures during training, facilitating their application to more diverse, uncurated data. I show that the internals of diffusion models can be used to meaningfully guide generation of new samples, without any further fine-tuning or supervision. Energy functions derived from a small set of primitive properties of denoiser activations can be combined to impose arbitrarily complex conditions on the iterative diffusion sampling procedure. This allows for control over attributes such as the position, shape, size, and appearance of any concept that can be described in text.I also demonstrate that the distribution learned by a text-to-image model can be distilled to generate compositional 3D scenes. Predominant approaches focus on creating 3D objects in isolation rather than scenes with several entities interacting. I propose an architecture that, when optimized so its outputs are on-manifold for the image generator, creates 3D scenes decomposed into the objects they contain. This provides evidence that scale alone suffices for a model to infer the actual 3D structure latent to a world it observes only through 2D images.Finally, I conclude with a perspective on the interplay between emergence, control, interpretability, and scale, and humbly attempt to relate these themes to the pursuit of intelligence.
일반주제명  
Computer science
일반주제명  
Engineering
일반주제명  
Information technology
키워드  
Generative adversarial network
키워드  
Bottlenecks
키워드  
Visual data
키워드  
Text-to-image model
키워드  
3D scenes
기타저자  
University of California, Berkeley Computer Science
기본자료저록  
Dissertations Abstracts International. 86-04B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017161672
■00520250211151428
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798384447214
■035    ▼a(MiAaPQ)AAI31295008
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aEpstein,  Dave.
■24510▼aDisentangled  Visual  Generative  Models
■260    ▼a[Sl]▼bUniversity  of  California,  Berkeley▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a116  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-04,  Section:  B.
■500    ▼aAdvisor:  Efros,  Alexei  A.
■5021  ▼aThesis  (Ph.D.)--University  of  California,  Berkeley,  2024.
■520    ▼aGenerative  modeling  promises  an  elegant  solution  to  learning  about  high-dimensional  data  distributions  such  as  images  and  videos  -  but  how  can  we  expose  and  utilize  the  rich  structure  these  models  discover?  Rather  than  just  drawing  new  samples,  how  can  an  agent  actually  harness  p(x)  as  a  source  of  knowledge  about  how  our  world  works?  This  thesis  explores  scalable  inductive  biases  that  unlock  a  generative  model's  disentangled  understanding  of  visual  data,  enabling  much  richer  interaction  and  control  as  a  result.First,  I  propose  a  representation  of  scenes  as  collections  of  feature  "blobs",  where  a  generative  adversarial  network  (GAN)  learns  -  without  any  labels  -  to  bind  each  blob  to  a  different  object  in  the  images  it  creates.  This  allows  GANs  to  more  gracefully  model  compositional  scenes,  in  contrast  to  typical  unconditional  models  which  are  constrained  to  highly-aligned  single-object  data.  The  trained  model's  representation  can  easily  be  modified  to  counterfactually  manipulate  objects  in  both  generated  and  real  images. Next,  I  consider  methods  that  do  not  impose  bottlenecks  on  architectures  during  training,  facilitating  their  application  to  more  diverse,  uncurated  data.  I  show  that  the  internals  of  diffusion  models  can  be  used  to  meaningfully  guide  generation  of  new  samples,  without  any  further  fine-tuning  or  supervision.  Energy  functions  derived  from  a  small  set  of  primitive  properties  of  denoiser  activations  can  be  combined  to  impose  arbitrarily  complex  conditions  on  the  iterative  diffusion  sampling  procedure.  This  allows  for  control  over  attributes  such  as  the  position,  shape,  size,  and  appearance  of  any  concept  that  can  be  described  in  text.I  also  demonstrate  that  the  distribution  learned  by  a  text-to-image  model  can  be  distilled  to  generate  compositional  3D  scenes.  Predominant  approaches  focus  on  creating  3D  objects  in  isolation  rather  than  scenes  with  several  entities  interacting.  I  propose  an  architecture  that,  when  optimized  so  its  outputs  are  on-manifold  for  the  image  generator,  creates  3D  scenes  decomposed  into  the  objects  they  contain.  This  provides  evidence  that  scale  alone  suffices  for  a  model  to  infer  the  actual  3D  structure  latent  to  a  world  it  observes  only  through  2D  images.Finally,  I  conclude  with  a  perspective  on  the  interplay  between  emergence,  control,  interpretability,  and  scale,  and  humbly  attempt  to  relate  these  themes  to  the  pursuit  of  intelligence.
■590    ▼aSchool  code:  0028.
■650  4▼aComputer  science
■650  4▼aEngineering
■650  4▼aInformation  technology
■653    ▼aGenerative  adversarial  network
■653    ▼aBottlenecks
■653    ▼aVisual  data
■653    ▼aText-to-image  model
■653    ▼a3D  scenes
■690    ▼a0984
■690    ▼a0489
■690    ▼a0800
■690    ▼a0537
■71020▼aUniversity  of  California,  Berkeley▼bComputer  Science.
■7730  ▼tDissertations  Abstracts  International▼g86-04B.
■790    ▼a0028
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17161672▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

Preview

Export

ChatGPT Discussion

AI Recommended Related Books


    New Books MORE
    Statistics for the past 3 years. Go to brief

    Подробнее информация.

    • Бронирование
    • не существует
    • моя папка
    • Первый запрос зрения
    • Non-Book Loan Application
    • Nighttime Book Loan Application
    материал
    Reg No. Количество платежных Местоположение статус Ленд информации
    TF13543 전자도서 대출가능 My Folder 부재도서신고 비도서대출신청 야간 도서대출신청

    * Бронирование доступны в заимствований книги. Чтобы сделать предварительный заказ, пожалуйста, нажмите кнопку бронирование

    Books borrowed together with this book

    Related Popular Books

    Available after logging in.