서브메뉴
검색
Generative Deep Learning: Towards Better Visual Representations and Multimodal
Generative Deep Learning: Towards Better Visual Representations and Multimodal
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260209102848
- ISBN
- 9798291563649
- DDC
- 004
- 저자명
- Xu, Xingqian.
- 서명/저자
- Generative Deep Learning: Towards Better Visual Representations and Multimodal
- 발행사항
- [Sl] : University of Illinois at Urbana-Champaign, 2023
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2023
- 형태사항
- 125 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-02, Section: B.
- 주기사항
- Advisor: Shi, Humphrey.
- 학위논문주기
- Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2023.
- 초록/해제
- 요약Generative AI aims to formulate certain types of data distribution so that it can generate new data instances mimicking true samples from the underlining distribution. It is also worth mentioning that in Computer Vision, generative and discriminative models are two major categories. While the latter aims to accurately predict classes, object locations, segmentations, etc. based on a specific data instance, the former explores and manufactures the complex data manifold. One may argue that generative AI in computer vision needs to be more advanced due to its intentions to simulate real-world data in unrestricted domains with tremendous complexity. Yet, even with the most complex network design, formulating the exact data distribution in our natural world is most likely inconceivable, thus leaving much space to improve.With the recent technology booms in Generative AI, nowadays researchers and engineers create high-performing generative solutions that start to handle real-world needs as commercial products, and luckily this thesis takes part as well. In this thesis, the author targets to further push generative AI performance by exploring the best possible visual representation form (i.e. neural implicit embedding, spectral-domain representation, transformer-based representation) that captures as much visual information as possible. Unquestionably, data representation is a critical premise to Generative AI as it reveals the upper bound of the model's capacity. Moreover, from a broader but less precise angle, the goal of generative modeling, simulating accurate data distribution, also serves as a kind of representation learning. In the final part of this thesis, the author also investigated the topic that goes beyond visual representation, towards more general forms of cross-modal representations that fit multiple types of data modalities, which is a heuristic step towards an even more challenging quest: General AI.This thesis begins with UltraSR, which explores the implicit neural visual representation that well-fits image super-resolution by synthesizing image details with arbitrary upsampling scales. The core idea of UltraSR integrates implicit neural representation with learnable periodic encoding, formulating visual details in the high-frequency manifold in a continuous function. While UltraSR explores the neural visual representation, Spectral Hint GAN (SH-GAN) takes a diverse routine that deeply involves visual features in the frequency domain for image completion. SH-GAN proposed a novel spectral network module: Spectral Hint Unit (SHU), together with two new strategies: Heterogeneous Filtering and Gaussian Split. SH-GAN outperforms prior image completion methods due to the followings: effective inpaint low-frequency image structure via StyleGAN-based co-modulation framework, and effective inpaint high-frequency image texture via SHU. The recent process in text-to-image (T2I) diffusion models inspires us to explore a new work Prompt-Free Diffusion, in which we substitute CLIP text encoder with SeeCoder to capture visual cues, removing the need for prompts from the T2I system. SeeCoder automatically distillates all sorts of visual cues, including but not limited to semantics, texture, backgrounds, etc., and transpassing them to diffusion models. Our synthetic results are both high-quality and closely follow the reference visual cues encoded by SeeCoder. Parallel with Prompt-Free Diffusion, we proposed Versatile Diffusion, which is the first work proposing a unified multimodal multi-flow diffusion pipeline that uniformly handles numerous cross-modal tasks, generating images, text, and variations. Versatile Diffusion has a broader scope in which our goal is to combine the representation of different modalities into one generative network, making a bold step toward universal generative AI.In conclusion, all works have provided valuable insights into data representation, among which UltraSR, SH-GAN, and Prompt-Free Diffusion actively explore the best visual representation under three schemes: implicit neural representation, spectral-domain representation, and transformer-based representation. In the last part, Versatile Diffusion explores a unified representation and generation of image, text, and image-text cross-modals. UltraSR outperforms the baseline model by 0.05 dB on DIV2K across all scales. SHGAN reaches FID 3.41 on dataset FFHQ and 7.10 on dataset Places2, acquiring a new state-of-the-art in the large-scale free-form image completion task. Prompt-Free Diffusion and SeeCoder fulfill the popular exemplar-based image generation task with stunning quality. Versatile Diffusion's CLIP-similarity are 0.269 and 0.858; FIDs are 11.20 and 4.57 on Coco2014, measuring Text-to-Image and Image-Variation, outperforms the baseline Stable Diffusion in all aspects.
- 일반주제명
- Computer science
- 일반주제명
- Electrical engineering
- 일반주제명
- Computer engineering
- 키워드
- Generative model
- 키워드
- Computer vision
- 키워드
- Deep learning
- 기타저자
- University of Illinois at Urbana-Champaign Electrical & Computer Eng
- 기본자료저록
- Dissertations Abstracts International. 87-02B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260203s2023 us c eng d■001000017365886
■00520260209102848
■006m o d
■007cr#unu||||||||
■020 ▼a9798291563649
■035 ▼a(MiAaPQ)AAI32271347
■035 ▼a(MiAaPQ)httphdlhandlenet2142121491
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aXu, Xingqian.
■24510▼aGenerative Deep Learning: Towards Better Visual Representations and Multimodal
■260 ▼a[Sl]▼bUniversity of Illinois at Urbana-Champaign▼c2023
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2023
■300 ▼a125 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-02, Section: B.
■500 ▼aAdvisor: Shi, Humphrey.
■5021 ▼aThesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2023.
■520 ▼aGenerative AI aims to formulate certain types of data distribution so that it can generate new data instances mimicking true samples from the underlining distribution. It is also worth mentioning that in Computer Vision, generative and discriminative models are two major categories. While the latter aims to accurately predict classes, object locations, segmentations, etc. based on a specific data instance, the former explores and manufactures the complex data manifold. One may argue that generative AI in computer vision needs to be more advanced due to its intentions to simulate real-world data in unrestricted domains with tremendous complexity. Yet, even with the most complex network design, formulating the exact data distribution in our natural world is most likely inconceivable, thus leaving much space to improve.With the recent technology booms in Generative AI, nowadays researchers and engineers create high-performing generative solutions that start to handle real-world needs as commercial products, and luckily this thesis takes part as well. In this thesis, the author targets to further push generative AI performance by exploring the best possible visual representation form (i.e. neural implicit embedding, spectral-domain representation, transformer-based representation) that captures as much visual information as possible. Unquestionably, data representation is a critical premise to Generative AI as it reveals the upper bound of the model's capacity. Moreover, from a broader but less precise angle, the goal of generative modeling, simulating accurate data distribution, also serves as a kind of representation learning. In the final part of this thesis, the author also investigated the topic that goes beyond visual representation, towards more general forms of cross-modal representations that fit multiple types of data modalities, which is a heuristic step towards an even more challenging quest: General AI.This thesis begins with UltraSR, which explores the implicit neural visual representation that well-fits image super-resolution by synthesizing image details with arbitrary upsampling scales. The core idea of UltraSR integrates implicit neural representation with learnable periodic encoding, formulating visual details in the high-frequency manifold in a continuous function. While UltraSR explores the neural visual representation, Spectral Hint GAN (SH-GAN) takes a diverse routine that deeply involves visual features in the frequency domain for image completion. SH-GAN proposed a novel spectral network module: Spectral Hint Unit (SHU), together with two new strategies: Heterogeneous Filtering and Gaussian Split. SH-GAN outperforms prior image completion methods due to the followings: effective inpaint low-frequency image structure via StyleGAN-based co-modulation framework, and effective inpaint high-frequency image texture via SHU. The recent process in text-to-image (T2I) diffusion models inspires us to explore a new work Prompt-Free Diffusion, in which we substitute CLIP text encoder with SeeCoder to capture visual cues, removing the need for prompts from the T2I system. SeeCoder automatically distillates all sorts of visual cues, including but not limited to semantics, texture, backgrounds, etc., and transpassing them to diffusion models. Our synthetic results are both high-quality and closely follow the reference visual cues encoded by SeeCoder. Parallel with Prompt-Free Diffusion, we proposed Versatile Diffusion, which is the first work proposing a unified multimodal multi-flow diffusion pipeline that uniformly handles numerous cross-modal tasks, generating images, text, and variations. Versatile Diffusion has a broader scope in which our goal is to combine the representation of different modalities into one generative network, making a bold step toward universal generative AI.In conclusion, all works have provided valuable insights into data representation, among which UltraSR, SH-GAN, and Prompt-Free Diffusion actively explore the best visual representation under three schemes: implicit neural representation, spectral-domain representation, and transformer-based representation. In the last part, Versatile Diffusion explores a unified representation and generation of image, text, and image-text cross-modals. UltraSR outperforms the baseline model by 0.05 dB on DIV2K across all scales. SHGAN reaches FID 3.41 on dataset FFHQ and 7.10 on dataset Places2, acquiring a new state-of-the-art in the large-scale free-form image completion task. Prompt-Free Diffusion and SeeCoder fulfill the popular exemplar-based image generation task with stunning quality. Versatile Diffusion's CLIP-similarity are 0.269 and 0.858; FIDs are 11.20 and 4.57 on Coco2014, measuring Text-to-Image and Image-Variation, outperforms the baseline Stable Diffusion in all aspects.
■590 ▼aSchool code: 0090.
■650 4▼aComputer science
■650 4▼aElectrical engineering
■650 4▼aComputer engineering
■653 ▼aGenerative model
■653 ▼aRepresentation learning
■653 ▼aMultimodal learning
■653 ▼aComputer vision
■653 ▼aDeep learning
■690 ▼a0984
■690 ▼a0544
■690 ▼a0464
■71020▼aUniversity of Illinois at Urbana-Champaign▼bElectrical & Computer Eng.
■7730 ▼tDissertations Abstracts International▼g87-02B.
■790 ▼a0090
■791 ▼aPh.D.
■792 ▼a2023
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17365886▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


