서브메뉴
검색
Learning and Optimal Control of Dynamic Stochastic Systems: Robustness and Scalability
Learning and Optimal Control of Dynamic Stochastic Systems: Robustness and Scalability
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202104856
- ISBN
- 9798288816406
- DDC
- 519.7
- 저자명
- Wang, Shengbo.
- 서명/저자
- Learning and Optimal Control of Dynamic Stochastic Systems: Robustness and Scalability
- 발행사항
- [Sl] : Stanford University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 295 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-02, Section: B.
- 주기사항
- Advisor: Mancilla, Jose Blanchet;Glynn, Peter.
- 학위논문주기
- Thesis (Ph.D.)--Stanford University, 2025.
- 초록/해제
- 요약This thesis develops principled methodologies for the learning and optimal control of dynamic stochastic systems. It addresses two central challenges in modern decision-making: ensuring reliable performance under model misspecification or distributional shifts between training and deployment, and enabling scalable learning and control in large, complex systems. To this end, the thesis introduces new modeling frameworks, algorithmic designs, and theoretical analyses that advance both the statistical and computational frontiers of reinforcement learning and stochastic control, achieving robust and scalable policy learning.Chapters 2-4 focus on three complementary aspects of distributionally robust (DR) policy learning. Chapter 2 establishes a general framework of DR Markov decision processes (MDPs), promoting robustness by requiring the controller to adapt to worst-case deviations in the system evolution dynamics. It provides a complete characterization of when dynamic programming principles hold under various modeling assumptions, guided by information-adaptivity and geometric considerations. Chapter 3 leverages simulation-based methods to design computation- and memory-efficient, model-free DR Q-learning algorithms, including a variance-reduced variant, that avoid full model estimation and achieve near-optimal sample complexity. Chapter 4 extends the DR policy learning paradigm to continuous-state systems. It first establishes dynamic programming equations for the DR stochastic control formulation, then develops a learning paradigm that enables uniform estimation of the DR value function at parametric rates, even in the absence of parametric assumptions on the underlying data. This result is shown to be minimax-optimal, and the parametric rate exemplifies a statistically scalable learning paradigm.Chapter 5 resolves a longstanding open problem in average-reward reinforcement learning by proposing the first algorithm that achieves a sample complexity matching the theoretical lower bound for uniformly ergodic MDPs. The work shows that stability, captured via mixing properties of the MDP, can be systematically leveraged to reduce the fundamental complexity of learning. By exploiting this structure, the algorithm achieves statistically optimal performance, demonstrating that stability is not just desirable for control, but essential for efficient learning. This insight further reinforces the role of system-specific structures in enabling statistical scalability. Finally, Chapter 6 turns to scalability in overparameterized systems, where the use of expressive neural network models often introduces significant computational bottlenecks. Focusing on the optimization of a value functional associated with an overparameterized stochastic differential equation (SDE) with jumps, it introduces an unbiased gradient estimator whose simulation cost remains insensitive to increasingly large parameter dimensions. This innovation enables efficient optimization in high-dimensional environments, where traditional gradient methods incur prohibitive computation times. Applications include neural SDEs, stochastic control, and simulation-based policy learning in large-scale systems.Together, these contributions form a principled and well-structured approach to robust and scalable learning in dynamic stochastic environments. They yield practical tools for reliable decision-making in both tabular and continuous settings and advance our understanding of statistical and computational scalability in data-driven control.
- 일반주제명
- Dynamic programming
- 일반주제명
- Engineering
- 기타저자
- Stanford University.
- 기본자료저록
- Dissertations Abstracts International. 87-02B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017359253
■00520260202104856
■006m o d
■007cr#unu||||||||
■020 ▼a9798288816406
■035 ▼a(MiAaPQ)AAI32201012
■035 ▼a(MiAaPQ)Stanfordwp673fm8582
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a519.7
■1001 ▼aWang, Shengbo.
■24510▼aLearning and Optimal Control of Dynamic Stochastic Systems: Robustness and Scalability
■260 ▼a[Sl]▼bStanford University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a295 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-02, Section: B.
■500 ▼aAdvisor: Mancilla, Jose Blanchet;Glynn, Peter.
■5021 ▼aThesis (Ph.D.)--Stanford University, 2025.
■520 ▼aThis thesis develops principled methodologies for the learning and optimal control of dynamic stochastic systems. It addresses two central challenges in modern decision-making: ensuring reliable performance under model misspecification or distributional shifts between training and deployment, and enabling scalable learning and control in large, complex systems. To this end, the thesis introduces new modeling frameworks, algorithmic designs, and theoretical analyses that advance both the statistical and computational frontiers of reinforcement learning and stochastic control, achieving robust and scalable policy learning.Chapters 2-4 focus on three complementary aspects of distributionally robust (DR) policy learning. Chapter 2 establishes a general framework of DR Markov decision processes (MDPs), promoting robustness by requiring the controller to adapt to worst-case deviations in the system evolution dynamics. It provides a complete characterization of when dynamic programming principles hold under various modeling assumptions, guided by information-adaptivity and geometric considerations. Chapter 3 leverages simulation-based methods to design computation- and memory-efficient, model-free DR Q-learning algorithms, including a variance-reduced variant, that avoid full model estimation and achieve near-optimal sample complexity. Chapter 4 extends the DR policy learning paradigm to continuous-state systems. It first establishes dynamic programming equations for the DR stochastic control formulation, then develops a learning paradigm that enables uniform estimation of the DR value function at parametric rates, even in the absence of parametric assumptions on the underlying data. This result is shown to be minimax-optimal, and the parametric rate exemplifies a statistically scalable learning paradigm.Chapter 5 resolves a longstanding open problem in average-reward reinforcement learning by proposing the first algorithm that achieves a sample complexity matching the theoretical lower bound for uniformly ergodic MDPs. The work shows that stability, captured via mixing properties of the MDP, can be systematically leveraged to reduce the fundamental complexity of learning. By exploiting this structure, the algorithm achieves statistically optimal performance, demonstrating that stability is not just desirable for control, but essential for efficient learning. This insight further reinforces the role of system-specific structures in enabling statistical scalability. Finally, Chapter 6 turns to scalability in overparameterized systems, where the use of expressive neural network models often introduces significant computational bottlenecks. Focusing on the optimization of a value functional associated with an overparameterized stochastic differential equation (SDE) with jumps, it introduces an unbiased gradient estimator whose simulation cost remains insensitive to increasingly large parameter dimensions. This innovation enables efficient optimization in high-dimensional environments, where traditional gradient methods incur prohibitive computation times. Applications include neural SDEs, stochastic control, and simulation-based policy learning in large-scale systems.Together, these contributions form a principled and well-structured approach to robust and scalable learning in dynamic stochastic environments. They yield practical tools for reliable decision-making in both tabular and continuous settings and advance our understanding of statistical and computational scalability in data-driven control.
■590 ▼aSchool code: 0212.
■650 4▼aDynamic programming
■650 4▼aEngineering
■653 ▼aDistributionally robust policy learning
■653 ▼aMarkov decision processes
■690 ▼a0454
■690 ▼a0537
■71020▼aStanford University.
■7730 ▼tDissertations Abstracts International▼g87-02B.
■790 ▼a0212
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17359253▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


