본문

서브메뉴

Learning and Optimal Control of Dynamic Stochastic Systems: Robustness and Scalability
Learning and Optimal Control of Dynamic Stochastic Systems: Robustness and Scalability
Learning and Optimal Control of Dynamic Stochastic Systems: Robustness and Scalability

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202104856
ISBN  
9798288816406
DDC  
519.7
저자명  
Wang, Shengbo.
서명/저자  
Learning and Optimal Control of Dynamic Stochastic Systems: Robustness and Scalability
발행사항  
[Sl] : Stanford University, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
295 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-02, Section: B.
주기사항  
Advisor: Mancilla, Jose Blanchet;Glynn, Peter.
학위논문주기  
Thesis (Ph.D.)--Stanford University, 2025.
초록/해제  
요약This thesis develops principled methodologies for the learning and optimal control of dynamic stochastic systems. It addresses two central challenges in modern decision-making: ensuring reliable performance under model misspecification or distributional shifts between training and deployment, and enabling scalable learning and control in large, complex systems. To this end, the thesis introduces new modeling frameworks, algorithmic designs, and theoretical analyses that advance both the statistical and computational frontiers of reinforcement learning and stochastic control, achieving robust and scalable policy learning.Chapters 2-4 focus on three complementary aspects of distributionally robust (DR) policy learning. Chapter 2 establishes a general framework of DR Markov decision processes (MDPs), promoting robustness by requiring the controller to adapt to worst-case deviations in the system evolution dynamics. It provides a complete characterization of when dynamic programming principles hold under various modeling assumptions, guided by information-adaptivity and geometric considerations. Chapter 3 leverages simulation-based methods to design computation- and memory-efficient, model-free DR Q-learning algorithms, including a variance-reduced variant, that avoid full model estimation and achieve near-optimal sample complexity. Chapter 4 extends the DR policy learning paradigm to continuous-state systems. It first establishes dynamic programming equations for the DR stochastic control formulation, then develops a learning paradigm that enables uniform estimation of the DR value function at parametric rates, even in the absence of parametric assumptions on the underlying data. This result is shown to be minimax-optimal, and the parametric rate exemplifies a statistically scalable learning paradigm.Chapter 5 resolves a longstanding open problem in average-reward reinforcement learning by proposing the first algorithm that achieves a sample complexity matching the theoretical lower bound for uniformly ergodic MDPs. The work shows that stability, captured via mixing properties of the MDP, can be systematically leveraged to reduce the fundamental complexity of learning. By exploiting this structure, the algorithm achieves statistically optimal performance, demonstrating that stability is not just desirable for control, but essential for efficient learning. This insight further reinforces the role of system-specific structures in enabling statistical scalability. Finally, Chapter 6 turns to scalability in overparameterized systems, where the use of expressive neural network models often introduces significant computational bottlenecks. Focusing on the optimization of a value functional associated with an overparameterized stochastic differential equation (SDE) with jumps, it introduces an unbiased gradient estimator whose simulation cost remains insensitive to increasingly large parameter dimensions. This innovation enables efficient optimization in high-dimensional environments, where traditional gradient methods incur prohibitive computation times. Applications include neural SDEs, stochastic control, and simulation-based policy learning in large-scale systems.Together, these contributions form a principled and well-structured approach to robust and scalable learning in dynamic stochastic environments. They yield practical tools for reliable decision-making in both tabular and continuous settings and advance our understanding of statistical and computational scalability in data-driven control. 
일반주제명  
Dynamic programming
일반주제명  
Engineering
키워드  
Distributionally robust policy learning
키워드  
Markov decision processes
기타저자  
Stanford University.
기본자료저록  
Dissertations Abstracts International. 87-02B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017359253
■00520260202104856
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798288816406
■035    ▼a(MiAaPQ)AAI32201012
■035    ▼a(MiAaPQ)Stanfordwp673fm8582
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a519.7
■1001  ▼aWang,  Shengbo.
■24510▼aLearning  and  Optimal  Control  of  Dynamic  Stochastic  Systems:  Robustness  and  Scalability
■260    ▼a[Sl]▼bStanford  University▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a295  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-02,  Section:  B.
■500    ▼aAdvisor:  Mancilla,  Jose  Blanchet;Glynn,  Peter.
■5021  ▼aThesis  (Ph.D.)--Stanford  University,  2025.
■520    ▼aThis  thesis  develops  principled  methodologies  for  the  learning  and  optimal  control  of  dynamic  stochastic  systems.  It  addresses  two  central  challenges  in  modern  decision-making:  ensuring  reliable  performance  under  model  misspecification  or  distributional  shifts  between  training  and  deployment,  and  enabling  scalable  learning  and  control  in  large,  complex  systems.  To  this  end,  the  thesis  introduces  new  modeling  frameworks,  algorithmic  designs,  and  theoretical  analyses  that  advance  both  the  statistical  and  computational  frontiers  of  reinforcement  learning  and  stochastic  control,  achieving  robust  and  scalable  policy  learning.Chapters  2-4  focus  on  three  complementary  aspects  of  distributionally  robust  (DR)  policy  learning.  Chapter  2  establishes  a  general  framework  of  DR  Markov  decision  processes  (MDPs),  promoting  robustness  by  requiring  the  controller  to  adapt  to  worst-case  deviations  in  the  system  evolution  dynamics.  It  provides  a  complete  characterization  of  when  dynamic  programming  principles  hold  under  various  modeling  assumptions,  guided  by  information-adaptivity  and  geometric  considerations.  Chapter  3  leverages  simulation-based  methods  to  design  computation-  and  memory-efficient,  model-free  DR  Q-learning  algorithms,  including  a  variance-reduced  variant,  that  avoid  full  model  estimation  and  achieve  near-optimal  sample  complexity.  Chapter  4  extends  the  DR  policy  learning  paradigm  to  continuous-state  systems.  It  first  establishes  dynamic  programming  equations  for  the  DR  stochastic  control  formulation,  then  develops  a  learning  paradigm  that  enables  uniform  estimation  of  the  DR  value  function  at  parametric  rates,  even  in  the  absence  of  parametric  assumptions  on  the  underlying  data.  This  result  is  shown  to  be  minimax-optimal,  and  the  parametric  rate  exemplifies  a  statistically  scalable  learning  paradigm.Chapter  5  resolves  a  longstanding  open  problem  in  average-reward  reinforcement  learning  by  proposing  the  first  algorithm  that  achieves  a  sample  complexity  matching  the  theoretical  lower  bound  for  uniformly  ergodic  MDPs.  The  work  shows  that  stability,  captured  via  mixing  properties  of  the  MDP,  can  be  systematically  leveraged  to  reduce  the  fundamental  complexity  of  learning.  By  exploiting  this  structure,  the  algorithm  achieves  statistically  optimal  performance,  demonstrating  that  stability  is  not  just  desirable  for  control,  but  essential  for  efficient  learning.  This  insight  further  reinforces  the  role  of  system-specific  structures  in  enabling  statistical  scalability. Finally,  Chapter  6  turns  to  scalability  in  overparameterized  systems,  where  the  use  of  expressive  neural  network  models  often  introduces  significant  computational  bottlenecks.  Focusing  on  the  optimization  of  a  value  functional  associated  with  an  overparameterized  stochastic  differential  equation  (SDE)  with  jumps,  it  introduces  an  unbiased  gradient  estimator  whose  simulation  cost  remains  insensitive  to  increasingly  large  parameter  dimensions.  This  innovation  enables  efficient  optimization  in  high-dimensional  environments,  where  traditional  gradient  methods  incur  prohibitive  computation  times.  Applications  include  neural  SDEs,  stochastic  control,  and  simulation-based  policy  learning  in  large-scale  systems.Together,  these  contributions  form  a  principled  and  well-structured  approach  to  robust  and  scalable  learning  in  dynamic  stochastic  environments.  They  yield  practical  tools  for  reliable  decision-making  in  both  tabular  and  continuous  settings  and  advance  our  understanding  of  statistical  and  computational  scalability  in  data-driven  control. 
■590    ▼aSchool  code:  0212.
■650  4▼aDynamic  programming
■650  4▼aEngineering
■653    ▼aDistributionally  robust  policy  learning
■653    ▼aMarkov  decision  processes
■690    ▼a0454
■690    ▼a0537
■71020▼aStanford  University.
■7730  ▼tDissertations  Abstracts  International▼g87-02B.
■790    ▼a0212
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17359253▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF15062 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.