서브메뉴
검색
Deep Generative Models for Video-Based Content Synthesis
Deep Generative Models for Video-Based Content Synthesis
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211151947
- ISBN
- 9798384015581
- DDC
- 621.3
- 저자명
- Wu, Xinyi.
- 서명/저자
- Deep Generative Models for Video-Based Content Synthesis
- 발행사항
- [Sl] : Northwestern University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 176 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-02, Section: A.
- 주기사항
- Advisor: Katsaggelos, Aggelos K.
- 학위논문주기
- Thesis (Ph.D.)--Northwestern University, 2024.
- 초록/해제
- 요약Video-based user-generated content (UGC) has increasingly become a vital component of our daily entertainment. Despite its popularity, creating high-quality videos remains a challenge for creators due to factors like labor costs, production expenses, and professional expertise. Deep generative models have emerged as a solution to assist in producing appealing UGC videos efficiently. Given the example of making animated movie videos, these models facilitate the entire production process, covering preparation, planning, and presentation. More specifically, they are capable of enriching the animation resource database, enabling automatic cinematography, and ensuring post-processing enhancement.In this work, we primarily focus on three critical aspects of automated UGC video production: actor pose synthesis, camera trajectory control, and video super-resolution. For the first aspect, we propose a harmony-aware human motion synthesis network that generates realistic actor poses synchronized to the given music. Differently from the previous methods, our approach not only tackles effective cross-modal transformation through meter-based GAN training but also maintains audio-visual consistency using a novel harmony loss driven by beat-level synchronization.To create an immersive viewer experience, we propose a camera control system to handle automatic cinematography following a well-recognized empirical rule-actor-camera synchronization. This system consists of two modules: aesthetic composition adjustment and camera trajectory synthesis. The former module optimizes aesthetic framing based on the regularization of the rule-of-thirds aesthetic principle and outputs an aesthetic camera placement for initialization. Subsequently, the latter module generates camera movements transferred from the actor's physical and psychological behaviors, managing spatial tracking and emotional styling. All these designs jointly contribute to a significant enhancement for the overall sense of immersion.In terms of the third aspect, we analyze video super-resolution (VSR), which allows us to enhance the resolution of the produced video at a low cost so as to further improve the viewer experience. Here, we propose a lightweight VSR model that is able to process 4K video in real time by incorporating the video quality assessment(VQA) strategy. Unlike traditional methods driven by spatial metrics like PSNR and SSIM, we integrate ST-RRED, a powerful VQA approach that separately measures spatial and temporal consistency aligning with human perception principles, into our loss functions. This guides us in reconstructing quality-aware perceptual features across space and time from low resolution to high resolution.We also provide results from quantitative and qualitative experiments to demonstrate the effectiveness of all our proposed generative models through their abilities to improve the production and generation quality of UGC videos.
- 일반주제명
- Electrical engineering
- 일반주제명
- Information science
- 기타저자
- Northwestern University Electrical Engineering
- 기본자료저록
- Dissertations Abstracts International. 86-02A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017162221
■00520250211151947
■006m o d
■007cr#unu||||||||
■020 ▼a9798384015581
■035 ▼a(MiAaPQ)AAI31327896
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a621.3
■1001 ▼aWu, Xinyi.▼0(orcid)0000-0003-0791-6305
■24510▼aDeep Generative Models for Video-Based Content Synthesis
■260 ▼a[Sl]▼bNorthwestern University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a176 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-02, Section: A.
■500 ▼aAdvisor: Katsaggelos, Aggelos K.
■5021 ▼aThesis (Ph.D.)--Northwestern University, 2024.
■520 ▼aVideo-based user-generated content (UGC) has increasingly become a vital component of our daily entertainment. Despite its popularity, creating high-quality videos remains a challenge for creators due to factors like labor costs, production expenses, and professional expertise. Deep generative models have emerged as a solution to assist in producing appealing UGC videos efficiently. Given the example of making animated movie videos, these models facilitate the entire production process, covering preparation, planning, and presentation. More specifically, they are capable of enriching the animation resource database, enabling automatic cinematography, and ensuring post-processing enhancement.In this work, we primarily focus on three critical aspects of automated UGC video production: actor pose synthesis, camera trajectory control, and video super-resolution. For the first aspect, we propose a harmony-aware human motion synthesis network that generates realistic actor poses synchronized to the given music. Differently from the previous methods, our approach not only tackles effective cross-modal transformation through meter-based GAN training but also maintains audio-visual consistency using a novel harmony loss driven by beat-level synchronization.To create an immersive viewer experience, we propose a camera control system to handle automatic cinematography following a well-recognized empirical rule-actor-camera synchronization. This system consists of two modules: aesthetic composition adjustment and camera trajectory synthesis. The former module optimizes aesthetic framing based on the regularization of the rule-of-thirds aesthetic principle and outputs an aesthetic camera placement for initialization. Subsequently, the latter module generates camera movements transferred from the actor's physical and psychological behaviors, managing spatial tracking and emotional styling. All these designs jointly contribute to a significant enhancement for the overall sense of immersion.In terms of the third aspect, we analyze video super-resolution (VSR), which allows us to enhance the resolution of the produced video at a low cost so as to further improve the viewer experience. Here, we propose a lightweight VSR model that is able to process 4K video in real time by incorporating the video quality assessment(VQA) strategy. Unlike traditional methods driven by spatial metrics like PSNR and SSIM, we integrate ST-RRED, a powerful VQA approach that separately measures spatial and temporal consistency aligning with human perception principles, into our loss functions. This guides us in reconstructing quality-aware perceptual features across space and time from low resolution to high resolution.We also provide results from quantitative and qualitative experiments to demonstrate the effectiveness of all our proposed generative models through their abilities to improve the production and generation quality of UGC videos.
■590 ▼aSchool code: 0163.
■650 4▼aElectrical engineering
■650 4▼aInformation science
■653 ▼aAudio-driven motion generation
■653 ▼aAutomatic camera trajectory control
■653 ▼aDeep generative models
■653 ▼aVideo super-resolution
■653 ▼aUser-generated content
■690 ▼a0544
■690 ▼a0723
■71020▼aNorthwestern University▼bElectrical Engineering.
■7730 ▼tDissertations Abstracts International▼g86-02A.
■790 ▼a0163
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17162221▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


