본문

서브메뉴

Generative Deep Learning: Towards Better Visual Representations and Multimodal
Generative Deep Learning: Towards Better Visual Representations and Multimodal
Generative Deep Learning: Towards Better Visual Representations and Multimodal

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260209102848
ISBN  
9798291563649
DDC  
004
저자명  
Xu, Xingqian.
서명/저자  
Generative Deep Learning: Towards Better Visual Representations and Multimodal
발행사항  
[Sl] : University of Illinois at Urbana-Champaign, 2023
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2023
형태사항  
125 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-02, Section: B.
주기사항  
Advisor: Shi, Humphrey.
학위논문주기  
Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2023.
초록/해제  
요약Generative AI aims to formulate certain types of data distribution so that it can generate new data instances mimicking true samples from the underlining distribution. It is also worth mentioning that in Computer Vision, generative and discriminative models are two major categories. While the latter aims to accurately predict classes, object locations, segmentations, etc. based on a specific data instance, the former explores and manufactures the complex data manifold. One may argue that generative AI in computer vision needs to be more advanced due to its intentions to simulate real-world data in unrestricted domains with tremendous complexity. Yet, even with the most complex network design, formulating the exact data distribution in our natural world is most likely inconceivable, thus leaving much space to improve.With the recent technology booms in Generative AI, nowadays researchers and engineers create high-performing generative solutions that start to handle real-world needs as commercial products, and luckily this thesis takes part as well. In this thesis, the author targets to further push generative AI performance by exploring the best possible visual representation form (i.e. neural implicit embedding, spectral-domain representation, transformer-based representation) that captures as much visual information as possible. Unquestionably, data representation is a critical premise to Generative AI as it reveals the upper bound of the model's capacity. Moreover, from a broader but less precise angle, the goal of generative modeling, simulating accurate data distribution, also serves as a kind of representation learning. In the final part of this thesis, the author also investigated the topic that goes beyond visual representation, towards more general forms of cross-modal representations that fit multiple types of data modalities, which is a heuristic step towards an even more challenging quest: General AI.This thesis begins with UltraSR, which explores the implicit neural visual representation that well-fits image super-resolution by synthesizing image details with arbitrary upsampling scales. The core idea of UltraSR integrates implicit neural representation with learnable periodic encoding, formulating visual details in the high-frequency manifold in a continuous function. While UltraSR explores the neural visual representation, Spectral Hint GAN (SH-GAN) takes a diverse routine that deeply involves visual features in the frequency domain for image completion. SH-GAN proposed a novel spectral network module: Spectral Hint Unit (SHU), together with two new strategies: Heterogeneous Filtering and Gaussian Split. SH-GAN outperforms prior image completion methods due to the followings: effective inpaint low-frequency image structure via StyleGAN-based co-modulation framework, and effective inpaint high-frequency image texture via SHU. The recent process in text-to-image (T2I) diffusion models inspires us to explore a new work Prompt-Free Diffusion, in which we substitute CLIP text encoder with SeeCoder to capture visual cues, removing the need for prompts from the T2I system. SeeCoder automatically distillates all sorts of visual cues, including but not limited to semantics, texture, backgrounds, etc., and transpassing them to diffusion models. Our synthetic results are both high-quality and closely follow the reference visual cues encoded by SeeCoder. Parallel with Prompt-Free Diffusion, we proposed Versatile Diffusion, which is the first work proposing a unified multimodal multi-flow diffusion pipeline that uniformly handles numerous cross-modal tasks, generating images, text, and variations. Versatile Diffusion has a broader scope in which our goal is to combine the representation of different modalities into one generative network, making a bold step toward universal generative AI.In conclusion, all works have provided valuable insights into data representation, among which UltraSR, SH-GAN, and Prompt-Free Diffusion actively explore the best visual representation under three schemes: implicit neural representation, spectral-domain representation, and transformer-based representation. In the last part, Versatile Diffusion explores a unified representation and generation of image, text, and image-text cross-modals. UltraSR outperforms the baseline model by 0.05 dB on DIV2K across all scales. SHGAN reaches FID 3.41 on dataset FFHQ and 7.10 on dataset Places2, acquiring a new state-of-the-art in the large-scale free-form image completion task. Prompt-Free Diffusion and SeeCoder fulfill the popular exemplar-based image generation task with stunning quality. Versatile Diffusion's CLIP-similarity are 0.269 and 0.858; FIDs are 11.20 and 4.57 on Coco2014, measuring Text-to-Image and Image-Variation, outperforms the baseline Stable Diffusion in all aspects.
일반주제명  
Computer science
일반주제명  
Electrical engineering
일반주제명  
Computer engineering
키워드  
Generative model
키워드  
Representation learning
키워드  
Multimodal learning
키워드  
Computer vision
키워드  
Deep learning
기타저자  
University of Illinois at Urbana-Champaign Electrical & Computer Eng
기본자료저록  
Dissertations Abstracts International. 87-02B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260203s2023        us                              c    eng  d
■001000017365886
■00520260209102848
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798291563649
■035    ▼a(MiAaPQ)AAI32271347
■035    ▼a(MiAaPQ)httphdlhandlenet2142121491
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aXu,  Xingqian.
■24510▼aGenerative  Deep  Learning:  Towards  Better  Visual  Representations  and  Multimodal
■260    ▼a[Sl]▼bUniversity  of  Illinois  at  Urbana-Champaign▼c2023
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2023
■300    ▼a125  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-02,  Section:  B.
■500    ▼aAdvisor:  Shi,  Humphrey.
■5021  ▼aThesis  (Ph.D.)--University  of  Illinois  at  Urbana-Champaign,  2023.
■520    ▼aGenerative  AI  aims  to  formulate  certain  types  of  data  distribution  so  that  it  can  generate  new  data  instances  mimicking  true  samples  from  the  underlining  distribution.  It  is  also  worth  mentioning  that  in  Computer  Vision,  generative  and  discriminative  models  are  two  major  categories.  While  the  latter  aims  to  accurately  predict  classes,  object  locations,  segmentations,  etc.  based  on  a  specific  data  instance,  the  former  explores  and  manufactures  the  complex  data  manifold.  One  may  argue  that  generative  AI  in  computer  vision  needs  to  be  more  advanced  due  to  its  intentions  to  simulate  real-world  data  in  unrestricted  domains  with  tremendous  complexity.  Yet,  even  with  the  most  complex  network  design,  formulating  the  exact  data  distribution  in  our  natural  world  is  most  likely  inconceivable,  thus  leaving  much  space  to  improve.With  the  recent  technology  booms  in  Generative  AI,  nowadays  researchers  and  engineers  create  high-performing  generative  solutions  that  start  to  handle  real-world  needs  as  commercial  products,  and  luckily  this  thesis  takes  part  as  well.  In  this  thesis,  the  author  targets  to  further  push  generative  AI  performance  by  exploring  the  best  possible  visual  representation  form  (i.e.  neural  implicit  embedding,  spectral-domain  representation,  transformer-based  representation)  that  captures  as  much  visual  information  as  possible.  Unquestionably,  data  representation  is  a  critical  premise  to  Generative  AI  as  it  reveals  the  upper  bound  of  the  model's  capacity.  Moreover,  from  a  broader  but  less  precise  angle,  the  goal  of  generative  modeling,  simulating  accurate  data  distribution,  also  serves  as  a  kind  of  representation  learning.  In  the  final  part  of  this  thesis,  the  author  also  investigated  the  topic  that  goes  beyond  visual  representation,  towards  more  general  forms  of  cross-modal  representations  that  fit  multiple  types  of  data  modalities,  which  is  a  heuristic  step  towards  an  even  more  challenging  quest:  General  AI.This  thesis  begins  with  UltraSR,  which  explores  the  implicit  neural  visual  representation  that  well-fits  image  super-resolution  by  synthesizing  image  details  with  arbitrary  upsampling  scales.  The  core  idea  of  UltraSR  integrates  implicit  neural  representation  with  learnable  periodic  encoding,  formulating  visual  details  in  the  high-frequency  manifold  in  a  continuous  function.  While  UltraSR  explores  the  neural  visual  representation,  Spectral  Hint  GAN  (SH-GAN)  takes  a  diverse  routine  that  deeply  involves  visual  features  in  the  frequency  domain  for  image  completion.  SH-GAN  proposed  a  novel  spectral  network  module:  Spectral  Hint  Unit  (SHU),  together  with  two  new  strategies:  Heterogeneous  Filtering  and  Gaussian  Split.  SH-GAN  outperforms  prior  image  completion  methods  due  to  the  followings:  effective  inpaint  low-frequency  image  structure  via  StyleGAN-based  co-modulation  framework,  and  effective  inpaint  high-frequency  image  texture  via  SHU.  The  recent  process  in  text-to-image  (T2I)  diffusion  models  inspires  us  to  explore  a  new  work  Prompt-Free  Diffusion,  in  which  we  substitute  CLIP  text  encoder  with  SeeCoder  to  capture  visual  cues,  removing  the  need  for  prompts  from  the  T2I  system.  SeeCoder  automatically  distillates  all  sorts  of  visual  cues,  including  but  not  limited  to  semantics,  texture,  backgrounds,  etc.,  and  transpassing  them  to  diffusion  models.  Our  synthetic  results  are  both  high-quality  and  closely  follow  the  reference  visual  cues  encoded  by  SeeCoder.  Parallel  with  Prompt-Free  Diffusion,  we  proposed  Versatile  Diffusion,  which  is  the  first  work  proposing  a  unified  multimodal  multi-flow  diffusion  pipeline  that  uniformly  handles  numerous  cross-modal  tasks,  generating  images,  text,  and  variations.  Versatile  Diffusion  has  a  broader  scope  in  which  our  goal  is  to  combine  the  representation  of  different  modalities  into  one  generative  network,  making  a  bold  step  toward  universal  generative  AI.In  conclusion,  all  works  have  provided  valuable  insights  into  data  representation,  among  which  UltraSR,  SH-GAN,  and  Prompt-Free  Diffusion  actively  explore  the  best  visual  representation  under  three  schemes:  implicit  neural  representation,  spectral-domain  representation,  and  transformer-based  representation.  In  the  last  part,  Versatile  Diffusion  explores  a  unified  representation  and  generation  of  image,  text,  and  image-text  cross-modals.  UltraSR  outperforms  the  baseline  model  by  0.05  dB  on  DIV2K  across  all  scales.  SHGAN  reaches  FID  3.41  on  dataset  FFHQ  and  7.10  on  dataset  Places2,  acquiring  a  new  state-of-the-art  in  the  large-scale  free-form  image  completion  task.  Prompt-Free  Diffusion  and  SeeCoder  fulfill  the  popular  exemplar-based  image  generation  task  with  stunning  quality.  Versatile  Diffusion's  CLIP-similarity  are  0.269  and  0.858;  FIDs  are  11.20  and  4.57  on  Coco2014,  measuring  Text-to-Image  and  Image-Variation,  outperforms  the  baseline  Stable  Diffusion  in  all  aspects.
■590    ▼aSchool  code:  0090.
■650  4▼aComputer  science
■650  4▼aElectrical  engineering
■650  4▼aComputer  engineering
■653    ▼aGenerative  model
■653    ▼aRepresentation  learning
■653    ▼aMultimodal  learning
■653    ▼aComputer  vision
■653    ▼aDeep  learning
■690    ▼a0984
■690    ▼a0544
■690    ▼a0464
■71020▼aUniversity  of  Illinois  at  Urbana-Champaign▼bElectrical  &  Computer  Eng.
■7730  ▼tDissertations  Abstracts  International▼g87-02B.
■790    ▼a0090
■791    ▼aPh.D.
■792    ▼a2023
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17365886▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF18992 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.