서브메뉴
검색
Stream Denoising and Efficient Diffusion Model Design for Real-Time Interactive Generation
Stream Denoising and Efficient Diffusion Model Design for Real-Time Interactive Generation
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105113
- ISBN
- 9798297601840
- DDC
- 629.8
- 저자명
- Kodaira, Akio.
- 서명/저자
- Stream Denoising and Efficient Diffusion Model Design for Real-Time Interactive Generation
- 발행사항
- [Sl] : University of California, Berkeley, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 101 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-04, Section: B.
- 주기사항
- Includes supplementary digital materials.
- 주기사항
- Advisor: Tomizuka, Masayoshi.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Berkeley, 2025.
- 초록/해제
- 요약Diffusion models have demonstrated exceptional capability in modeling complex, high-dimensional data, achieving state-of-the-art performance in tasks previously considered intractable, such as generating entirely novel images and videos. Recent developments, including Diffusion Policy for synthesizing robot action sequences, have achieved remarkable results, further elevating interest in the application of diffusion models to robotics.Despite these advances, most existing diffusion model applications and products remain oriented toward generating images in an offline manner, typically in response to user queries. Consequently, both research and practice have devoted limited attention to the dimensions of real-time performance and interactivity. Existing inference pipelines and model architectures are generally not optimized for such use cases, and the multi-step denoising process-operating over data chunks-renders conventional diffusion models ill-suited to interactive scenarios. For future adoption in robotics, it is essential to design architectures and inference pipelines that explicitly address real-time constraints and interactive operation.This dissertation introduces Stream Denoising, a novel denoising paradigm that enables diffusion models to perform inference in real time while maintaining interactive responsiveness.The first part of the dissertation presents StreamDiffusion, a training-free pipeline for real time operation of image generation models, followed by StreamV2V, an extension that improves long tail temporal consistency through mechanisms such as Feature Banks and Dynamic Merging.The second part addresses model training, introducing StreamDiT, a unified training method for transforming video-generation diffusion models into real-time, interactive video synthesis systems.Collectively, these contributions establish a unified set of concepts and methodologies for adapting diffusion models to real-time, interactive applications in the image and video domains, thereby bridging the gap between current generative modeling techniques and their future deployment in robotics.
- 일반주제명
- Robotics
- 일반주제명
- Computer engineering
- 일반주제명
- Mechanical engineering
- 키워드
- Diffusion models
- 키워드
- Stream Denoising
- 키워드
- Video generation
- 기타저자
- University of California, Berkeley Mechanical Engineering
- 기본자료저록
- Dissertations Abstracts International. 87-04B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017359394
■00520260202105113
■006m o d
■007cr#unu||||||||
■020 ▼a9798297601840
■035 ▼a(MiAaPQ)AAI32237216
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a629.8
■1001 ▼aKodaira, Akio.
■24510▼aStream Denoising and Efficient Diffusion Model Design for Real-Time Interactive Generation
■260 ▼a[Sl]▼bUniversity of California, Berkeley▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a101 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-04, Section: B.
■500 ▼aIncludes supplementary digital materials.
■500 ▼aAdvisor: Tomizuka, Masayoshi.
■5021 ▼aThesis (Ph.D.)--University of California, Berkeley, 2025.
■520 ▼aDiffusion models have demonstrated exceptional capability in modeling complex, high-dimensional data, achieving state-of-the-art performance in tasks previously considered intractable, such as generating entirely novel images and videos. Recent developments, including Diffusion Policy for synthesizing robot action sequences, have achieved remarkable results, further elevating interest in the application of diffusion models to robotics.Despite these advances, most existing diffusion model applications and products remain oriented toward generating images in an offline manner, typically in response to user queries. Consequently, both research and practice have devoted limited attention to the dimensions of real-time performance and interactivity. Existing inference pipelines and model architectures are generally not optimized for such use cases, and the multi-step denoising process-operating over data chunks-renders conventional diffusion models ill-suited to interactive scenarios. For future adoption in robotics, it is essential to design architectures and inference pipelines that explicitly address real-time constraints and interactive operation.This dissertation introduces Stream Denoising, a novel denoising paradigm that enables diffusion models to perform inference in real time while maintaining interactive responsiveness.The first part of the dissertation presents StreamDiffusion, a training-free pipeline for real time operation of image generation models, followed by StreamV2V, an extension that improves long tail temporal consistency through mechanisms such as Feature Banks and Dynamic Merging.The second part addresses model training, introducing StreamDiT, a unified training method for transforming video-generation diffusion models into real-time, interactive video synthesis systems.Collectively, these contributions establish a unified set of concepts and methodologies for adapting diffusion models to real-time, interactive applications in the image and video domains, thereby bridging the gap between current generative modeling techniques and their future deployment in robotics.
■590 ▼aSchool code: 0028.
■650 4▼aRobotics
■650 4▼aComputer engineering
■650 4▼aMechanical engineering
■653 ▼aAutoregressive diffusion
■653 ▼aDiffusion models
■653 ▼aInteractive scenarios
■653 ▼aReal-time constraints
■653 ▼aStream Denoising
■653 ▼aVideo generation
■690 ▼a0771
■690 ▼a0800
■690 ▼a0464
■690 ▼a0548
■71020▼aUniversity of California, Berkeley▼bMechanical Engineering.
■7730 ▼tDissertations Abstracts International▼g87-04B.
■790 ▼a0028
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17359394▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


