본문

서브메뉴

Stream Denoising and Efficient Diffusion Model Design for Real-Time Interactive Generation
Stream Denoising and Efficient Diffusion Model Design for Real-Time Interactive Generation
Stream Denoising and Efficient Diffusion Model Design for Real-Time Interactive Generation

Detailed Information

Material Type  
 단행본
 
0017359394
Date and Time of Latest Transaction  
20260202105113
ISBN  
9798297601840
DDC  
629.8
Author  
Kodaira, Akio.
Title/Author  
Stream Denoising and Efficient Diffusion Model Design for Real-Time Interactive Generation
Publish Info  
[Sl] : University of California, Berkeley, 2025
Publish Info  
Ann Arbor : ProQuest Dissertations & Theses, 2025
Material Info  
101 p
General Note  
Source: Dissertations Abstracts International, Volume: 87-04, Section: B.
General Note  
Includes supplementary digital materials.
General Note  
Advisor: Tomizuka, Masayoshi.
학위논문주기  
Thesis (Ph.D.)--University of California, Berkeley, 2025.
Abstracts/Etc  
요약Diffusion models have demonstrated exceptional capability in modeling complex, high-dimensional data, achieving state-of-the-art performance in tasks previously considered intractable, such as generating entirely novel images and videos. Recent developments, including Diffusion Policy for synthesizing robot action sequences, have achieved remarkable results, further elevating interest in the application of diffusion models to robotics.Despite these advances, most existing diffusion model applications and products remain oriented toward generating images in an offline manner, typically in response to user queries. Consequently, both research and practice have devoted limited attention to the dimensions of real-time performance and interactivity. Existing inference pipelines and model architectures are generally not optimized for such use cases, and the multi-step denoising process-operating over data chunks-renders conventional diffusion models ill-suited to interactive scenarios. For future adoption in robotics, it is essential to design architectures and inference pipelines that explicitly address real-time constraints and interactive operation.This dissertation introduces Stream Denoising, a novel denoising paradigm that enables diffusion models to perform inference in real time while maintaining interactive responsiveness.The first part of the dissertation presents StreamDiffusion, a training-free pipeline for real time operation of image generation models, followed by StreamV2V, an extension that improves long tail temporal consistency through mechanisms such as Feature Banks and Dynamic Merging.The second part addresses model training, introducing StreamDiT, a unified training method for transforming video-generation diffusion models into real-time, interactive video synthesis systems.Collectively, these contributions establish a unified set of concepts and methodologies for adapting diffusion models to real-time, interactive applications in the image and video domains, thereby bridging the gap between current generative modeling techniques and their future deployment in robotics.
Subject Added Entry-Topical Term  
Robotics
Subject Added Entry-Topical Term  
Computer engineering
Subject Added Entry-Topical Term  
Mechanical engineering
Index Term-Uncontrolled  
Autoregressive diffusion
Index Term-Uncontrolled  
Diffusion models
Index Term-Uncontrolled  
Interactive scenarios
Index Term-Uncontrolled  
Real-time constraints
Index Term-Uncontrolled  
Stream Denoising
Index Term-Uncontrolled  
Video generation
Added Entry-Corporate Name  
University of California, Berkeley Mechanical Engineering
Host Item Entry  
Dissertations Abstracts International. 87-04B.
Electronic Location and Access  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017359394
■00520260202105113
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798297601840
■035    ▼a(MiAaPQ)AAI32237216
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a629.8
■1001  ▼aKodaira,  Akio.
■24510▼aStream  Denoising  and  Efficient  Diffusion  Model  Design  for  Real-Time  Interactive  Generation
■260    ▼a[Sl]▼bUniversity  of  California,  Berkeley▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a101  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-04,  Section:  B.
■500    ▼aIncludes  supplementary  digital  materials.
■500    ▼aAdvisor:  Tomizuka,  Masayoshi.
■5021  ▼aThesis  (Ph.D.)--University  of  California,  Berkeley,  2025.
■520    ▼aDiffusion  models  have  demonstrated  exceptional  capability  in  modeling  complex,  high-dimensional  data,  achieving  state-of-the-art  performance  in  tasks  previously  considered  intractable,  such  as  generating  entirely  novel  images  and  videos.  Recent  developments,  including  Diffusion  Policy  for  synthesizing  robot  action  sequences,  have  achieved  remarkable  results,  further  elevating  interest  in  the  application  of  diffusion  models  to  robotics.Despite  these  advances,  most  existing  diffusion  model  applications  and  products  remain  oriented  toward  generating  images  in  an  offline  manner,  typically  in  response  to  user  queries.  Consequently,  both  research  and  practice  have  devoted  limited  attention  to  the  dimensions  of  real-time  performance  and  interactivity.  Existing  inference  pipelines  and  model  architectures  are  generally  not  optimized  for  such  use  cases,  and  the  multi-step  denoising  process-operating  over  data  chunks-renders  conventional  diffusion  models  ill-suited  to  interactive  scenarios.  For  future  adoption  in  robotics,  it  is  essential  to  design  architectures  and  inference  pipelines  that  explicitly  address  real-time  constraints  and  interactive  operation.This  dissertation  introduces  Stream  Denoising,  a  novel  denoising  paradigm  that  enables  diffusion  models  to  perform  inference  in  real  time  while  maintaining  interactive  responsiveness.The  first  part  of  the  dissertation  presents  StreamDiffusion,  a  training-free  pipeline  for  real  time  operation  of  image  generation  models,  followed  by  StreamV2V,  an  extension  that  improves  long  tail  temporal  consistency  through  mechanisms  such  as  Feature  Banks  and  Dynamic  Merging.The  second  part  addresses  model  training,  introducing  StreamDiT,  a  unified  training  method  for  transforming  video-generation  diffusion  models  into  real-time,  interactive  video  synthesis  systems.Collectively,  these  contributions  establish  a  unified  set  of  concepts  and  methodologies  for  adapting  diffusion  models  to  real-time,  interactive  applications  in  the  image  and  video  domains,  thereby  bridging  the  gap  between  current  generative  modeling  techniques  and  their  future  deployment  in  robotics.
■590    ▼aSchool  code:  0028.
■650  4▼aRobotics
■650  4▼aComputer  engineering
■650  4▼aMechanical  engineering
■653    ▼aAutoregressive  diffusion
■653    ▼aDiffusion  models
■653    ▼aInteractive  scenarios
■653    ▼aReal-time  constraints
■653    ▼aStream  Denoising
■653    ▼aVideo  generation
■690    ▼a0771
■690    ▼a0800
■690    ▼a0464
■690    ▼a0548
■71020▼aUniversity  of  California,  Berkeley▼bMechanical  Engineering.
■7730  ▼tDissertations  Abstracts  International▼g87-04B.
■790    ▼a0028
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17359394▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

Preview

Export

ChatGPT Discussion

AI Recommended Related Books


    New Books MORE
    Statistics for the past 3 years. Go to brief

    Detail Info.

    • Reservation
    • Not Exist
    • My Folder
    • First Request
    • Non-Book Loan Application
    • Nighttime Book Loan Application
    Material
    Reg No. Call No. Location Status Lend Info
    TF15854 전자도서 대출가능 My Folder 부재도서신고 비도서대출신청 야간 도서대출신청

    * Reservations are available in the borrowing book. To make reservations, Please click the reservation button

    Books borrowed together with this book

    Related Popular Books

    Available after logging in.