본문

서브메뉴

Deep Generative Models for Video-Based Content Synthesis
Deep Generative Models for Video-Based Content Synthesis
Deep Generative Models for Video-Based Content Synthesis

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211151947
ISBN  
9798384015581
DDC  
621.3
저자명  
Wu, Xinyi.
서명/저자  
Deep Generative Models for Video-Based Content Synthesis
발행사항  
[Sl] : Northwestern University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
176 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-02, Section: A.
주기사항  
Advisor: Katsaggelos, Aggelos K.
학위논문주기  
Thesis (Ph.D.)--Northwestern University, 2024.
초록/해제  
요약Video-based user-generated content (UGC) has increasingly become a vital component of our daily entertainment. Despite its popularity, creating high-quality videos remains a challenge for creators due to factors like labor costs, production expenses, and professional expertise. Deep generative models have emerged as a solution to assist in producing appealing UGC videos efficiently. Given the example of making animated movie videos, these models facilitate the entire production process, covering preparation, planning, and presentation. More specifically, they are capable of enriching the animation resource database, enabling automatic cinematography, and ensuring post-processing enhancement.In this work, we primarily focus on three critical aspects of automated UGC video production: actor pose synthesis, camera trajectory control, and video super-resolution. For the first aspect, we propose a harmony-aware human motion synthesis network that generates realistic actor poses synchronized to the given music. Differently from the previous methods, our approach not only tackles effective cross-modal transformation through meter-based GAN training but also maintains audio-visual consistency using a novel harmony loss driven by beat-level synchronization.To create an immersive viewer experience, we propose a camera control system to handle automatic cinematography following a well-recognized empirical rule-actor-camera synchronization. This system consists of two modules: aesthetic composition adjustment and camera trajectory synthesis. The former module optimizes aesthetic framing based on the regularization of the rule-of-thirds aesthetic principle and outputs an aesthetic camera placement for initialization. Subsequently, the latter module generates camera movements transferred from the actor's physical and psychological behaviors, managing spatial tracking and emotional styling. All these designs jointly contribute to a significant enhancement for the overall sense of immersion.In terms of the third aspect, we analyze video super-resolution (VSR), which allows us to enhance the resolution of the produced video at a low cost so as to further improve the viewer experience. Here, we propose a lightweight VSR model that is able to process 4K video in real time by incorporating the video quality assessment(VQA) strategy. Unlike traditional methods driven by spatial metrics like PSNR and SSIM, we integrate ST-RRED, a powerful VQA approach that separately measures spatial and temporal consistency aligning with human perception principles, into our loss functions. This guides us in reconstructing quality-aware perceptual features across space and time from low resolution to high resolution.We also provide results from quantitative and qualitative experiments to demonstrate the effectiveness of all our proposed generative models through their abilities to improve the production and generation quality of UGC videos.
일반주제명  
Electrical engineering
일반주제명  
Information science
키워드  
Audio-driven motion generation
키워드  
Automatic camera trajectory control
키워드  
Deep generative models
키워드  
Video super-resolution
키워드  
User-generated content
기타저자  
Northwestern University Electrical Engineering
기본자료저록  
Dissertations Abstracts International. 86-02A.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017162221
■00520250211151947
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798384015581
■035    ▼a(MiAaPQ)AAI31327896
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a621.3
■1001  ▼aWu,  Xinyi.▼0(orcid)0000-0003-0791-6305
■24510▼aDeep  Generative  Models  for  Video-Based  Content  Synthesis
■260    ▼a[Sl]▼bNorthwestern  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a176  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-02,  Section:  A.
■500    ▼aAdvisor:  Katsaggelos,  Aggelos  K.
■5021  ▼aThesis  (Ph.D.)--Northwestern  University,  2024.
■520    ▼aVideo-based  user-generated  content  (UGC)  has  increasingly  become  a  vital  component  of  our  daily  entertainment.  Despite  its  popularity,  creating  high-quality  videos  remains  a  challenge  for  creators  due  to  factors  like  labor  costs,  production  expenses,  and  professional  expertise.  Deep  generative  models  have  emerged  as  a  solution  to  assist  in  producing  appealing  UGC  videos  efficiently.  Given  the  example  of  making  animated  movie  videos,  these  models  facilitate  the  entire  production  process,  covering  preparation,  planning,  and  presentation.  More  specifically,  they  are  capable  of  enriching  the  animation  resource  database,  enabling  automatic  cinematography,  and  ensuring  post-processing  enhancement.In  this  work,  we  primarily  focus  on  three  critical  aspects  of  automated  UGC  video  production:  actor  pose  synthesis,  camera  trajectory  control,  and  video  super-resolution.  For  the  first  aspect,  we  propose  a  harmony-aware  human  motion  synthesis  network  that  generates  realistic  actor  poses  synchronized  to  the  given  music.  Differently  from  the  previous  methods,  our  approach  not  only  tackles  effective  cross-modal  transformation  through  meter-based  GAN  training  but  also  maintains  audio-visual  consistency  using  a  novel  harmony  loss  driven  by  beat-level  synchronization.To  create  an  immersive  viewer  experience,  we  propose  a  camera  control  system  to  handle  automatic  cinematography  following  a  well-recognized  empirical  rule-actor-camera  synchronization.  This  system  consists  of  two  modules:  aesthetic  composition  adjustment  and  camera  trajectory  synthesis.  The  former  module  optimizes  aesthetic  framing  based  on  the  regularization  of  the  rule-of-thirds  aesthetic  principle  and  outputs  an  aesthetic  camera  placement  for  initialization.  Subsequently,  the  latter  module  generates  camera  movements  transferred  from  the  actor's  physical  and  psychological  behaviors,  managing  spatial  tracking  and  emotional  styling.  All  these  designs  jointly  contribute  to  a  significant  enhancement  for  the  overall  sense  of  immersion.In  terms  of  the  third  aspect,  we  analyze  video  super-resolution  (VSR),  which  allows  us  to  enhance  the  resolution  of  the  produced  video  at  a  low  cost  so  as  to  further  improve  the  viewer  experience.  Here,  we  propose  a  lightweight  VSR  model  that  is  able  to  process  4K  video  in  real  time  by  incorporating  the  video  quality  assessment(VQA)  strategy.  Unlike  traditional  methods  driven  by  spatial  metrics  like  PSNR  and  SSIM,  we  integrate  ST-RRED,  a  powerful  VQA  approach  that  separately  measures  spatial  and  temporal  consistency  aligning  with  human  perception  principles,  into  our  loss  functions.  This  guides  us  in  reconstructing  quality-aware  perceptual  features  across  space  and  time  from  low  resolution  to  high  resolution.We  also  provide  results  from  quantitative  and  qualitative  experiments  to  demonstrate  the  effectiveness  of  all  our  proposed  generative  models  through  their  abilities  to  improve  the  production  and  generation  quality  of  UGC  videos.
■590    ▼aSchool  code:  0163.
■650  4▼aElectrical  engineering
■650  4▼aInformation  science
■653    ▼aAudio-driven  motion  generation
■653    ▼aAutomatic  camera  trajectory  control
■653    ▼aDeep  generative  models
■653    ▼aVideo  super-resolution
■653    ▼aUser-generated  content
■690    ▼a0544
■690    ▼a0723
■71020▼aNorthwestern  University▼bElectrical  Engineering.
■7730  ▼tDissertations  Abstracts  International▼g86-02A.
■790    ▼a0163
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17162221▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF13418 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.