서브메뉴
검색
On the Theory of Safe Reinforcement Learning
On the Theory of Safe Reinforcement Learning
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103129
- ISBN
- 9798288861796
- DDC
- 658
- 저자명
- Ying, Donghao.
- 서명/저자
- On the Theory of Safe Reinforcement Learning
- 발행사항
- [Sl] : University of California, Berkeley, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 169 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-01, Section: B.
- 주기사항
- Advisor: Lavaei, Javad.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Berkeley, 2025.
- 초록/해제
- 요약Reinforcement learning (RL) has achieved notable success across a wide range of artificial sequential decision-making tasks, consistently demonstrating strong empirical performance in controlled environments. However, applying RL techniques to complex, safety-critical real-world domains-such as autonomous driving, healthcare, and robotic control-remains a significant challenge. These settings often come with strict safety requirements, where violating constraints can have serious consequences. To bridge this gap, Safe Reinforcement Learning (Safe RL) has emerged as a promising approach aimed at enabling high-quality decision-making while rigorously enforcing predefined safety constraints. Despite its potential, Safe RL not only inherits the fundamental difficulties of standard RL-such as high sample complexity and exploration-exploitation trade-offs-but also introduces additional complexities. These include (1) the need to carefully balance performance objectives with safety requirements and (2) the challenge of ensuring algorithmic convergence under nonconcave objectives and nonconvex constraints.Motivated by these challenges, this thesis aims to advance the theoretical foundations and algorithmic tools for Safe RL, focusing on the framework of Constrained Markov Decision Processes (CMDPs). We develop new analytical insights and propose a set of methods tailored to three representative CMDP problems: (1) Under entropy regularization, we introduce a dual optimization framework that takes advantage of the smoothness induced by entropy regularization. This leads to improved convergence guarantees for both the optimality gap and constraint violations. (2) When the objective and constraints take more general forms than the standard CMDP, we show that despite losing cumulative structure, hidden concavity can be exploited to design policy-based algorithms with provable global convergence guarantees. (3) In the presence of multiple agents, we come up with a distributed primal-dual algorithm utilizing shadow rewards and communication truncation strategies, effectively addressing scalability issues while ensuring convergence to first-order stationary points.
- 일반주제명
- Industrial engineering
- 키워드
- Decision-making
- 기타저자
- University of California, Berkeley Industrial Engineering & Operations Research
- 기본자료저록
- Dissertations Abstracts International. 87-01B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017357087
■00520260202103129
■006m o d
■007cr#unu||||||||
■020 ▼a9798288861796
■035 ▼a(MiAaPQ)AAI31939249
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a658
■1001 ▼aYing, Donghao.
■24510▼aOn the Theory of Safe Reinforcement Learning
■260 ▼a[Sl]▼bUniversity of California, Berkeley▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a169 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-01, Section: B.
■500 ▼aAdvisor: Lavaei, Javad.
■5021 ▼aThesis (Ph.D.)--University of California, Berkeley, 2025.
■520 ▼aReinforcement learning (RL) has achieved notable success across a wide range of artificial sequential decision-making tasks, consistently demonstrating strong empirical performance in controlled environments. However, applying RL techniques to complex, safety-critical real-world domains-such as autonomous driving, healthcare, and robotic control-remains a significant challenge. These settings often come with strict safety requirements, where violating constraints can have serious consequences. To bridge this gap, Safe Reinforcement Learning (Safe RL) has emerged as a promising approach aimed at enabling high-quality decision-making while rigorously enforcing predefined safety constraints. Despite its potential, Safe RL not only inherits the fundamental difficulties of standard RL-such as high sample complexity and exploration-exploitation trade-offs-but also introduces additional complexities. These include (1) the need to carefully balance performance objectives with safety requirements and (2) the challenge of ensuring algorithmic convergence under nonconcave objectives and nonconvex constraints.Motivated by these challenges, this thesis aims to advance the theoretical foundations and algorithmic tools for Safe RL, focusing on the framework of Constrained Markov Decision Processes (CMDPs). We develop new analytical insights and propose a set of methods tailored to three representative CMDP problems: (1) Under entropy regularization, we introduce a dual optimization framework that takes advantage of the smoothness induced by entropy regularization. This leads to improved convergence guarantees for both the optimality gap and constraint violations. (2) When the objective and constraints take more general forms than the standard CMDP, we show that despite losing cumulative structure, hidden concavity can be exploited to design policy-based algorithms with provable global convergence guarantees. (3) In the presence of multiple agents, we come up with a distributed primal-dual algorithm utilizing shadow rewards and communication truncation strategies, effectively addressing scalability issues while ensuring convergence to first-order stationary points.
■590 ▼aSchool code: 0028.
■650 4▼aIndustrial engineering
■653 ▼aLagrangian duality
■653 ▼aSafe Reinforcement Learning
■653 ▼aReinforcement learning
■653 ▼aDecision-making
■653 ▼aAutonomous driving
■690 ▼a0796
■690 ▼a0800
■690 ▼a0546
■71020▼aUniversity of California, Berkeley▼bIndustrial Engineering & Operations Research.
■7730 ▼tDissertations Abstracts International▼g87-01B.
■790 ▼a0028
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357087▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


