본문

서브메뉴

On the Theory of Safe Reinforcement Learning
On the Theory of Safe Reinforcement Learning
On the Theory of Safe Reinforcement Learning

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202103129
ISBN  
9798288861796
DDC  
658
저자명  
Ying, Donghao.
서명/저자  
On the Theory of Safe Reinforcement Learning
발행사항  
[Sl] : University of California, Berkeley, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
169 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-01, Section: B.
주기사항  
Advisor: Lavaei, Javad.
학위논문주기  
Thesis (Ph.D.)--University of California, Berkeley, 2025.
초록/해제  
요약Reinforcement learning (RL) has achieved notable success across a wide range of artificial sequential decision-making tasks, consistently demonstrating strong empirical performance in controlled environments. However, applying RL techniques to complex, safety-critical real-world domains-such as autonomous driving, healthcare, and robotic control-remains a significant challenge. These settings often come with strict safety requirements, where violating constraints can have serious consequences. To bridge this gap, Safe Reinforcement Learning (Safe RL) has emerged as a promising approach aimed at enabling high-quality decision-making while rigorously enforcing predefined safety constraints. Despite its potential, Safe RL not only inherits the fundamental difficulties of standard RL-such as high sample complexity and exploration-exploitation trade-offs-but also introduces additional complexities. These include (1) the need to carefully balance performance objectives with safety requirements and (2) the challenge of ensuring algorithmic convergence under nonconcave objectives and nonconvex constraints.Motivated by these challenges, this thesis aims to advance the theoretical foundations and algorithmic tools for Safe RL, focusing on the framework of Constrained Markov Decision Processes (CMDPs). We develop new analytical insights and propose a set of methods tailored to three representative CMDP problems: (1) Under entropy regularization, we introduce a dual optimization framework that takes advantage of the smoothness induced by entropy regularization. This leads to improved convergence guarantees for both the optimality gap and constraint violations. (2) When the objective and constraints take more general forms than the standard CMDP, we show that despite losing cumulative structure, hidden concavity can be exploited to design policy-based algorithms with provable global convergence guarantees. (3) In the presence of multiple agents, we come up with a distributed primal-dual algorithm utilizing shadow rewards and communication truncation strategies, effectively addressing scalability issues while ensuring convergence to first-order stationary points.
일반주제명  
Industrial engineering
키워드  
Lagrangian duality
키워드  
Safe Reinforcement Learning
키워드  
Reinforcement learning
키워드  
Decision-making
키워드  
Autonomous driving
기타저자  
University of California, Berkeley Industrial Engineering & Operations Research
기본자료저록  
Dissertations Abstracts International. 87-01B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017357087
■00520260202103129
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798288861796
■035    ▼a(MiAaPQ)AAI31939249
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a658
■1001  ▼aYing,  Donghao.
■24510▼aOn  the  Theory  of  Safe  Reinforcement  Learning
■260    ▼a[Sl]▼bUniversity  of  California,  Berkeley▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a169  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-01,  Section:  B.
■500    ▼aAdvisor:  Lavaei,  Javad.
■5021  ▼aThesis  (Ph.D.)--University  of  California,  Berkeley,  2025.
■520    ▼aReinforcement  learning  (RL)  has  achieved  notable  success  across  a  wide  range  of  artificial  sequential  decision-making  tasks,  consistently  demonstrating  strong  empirical  performance  in  controlled  environments.  However,  applying  RL  techniques  to  complex,  safety-critical  real-world  domains-such  as  autonomous  driving,  healthcare,  and  robotic  control-remains  a  significant  challenge.  These  settings  often  come  with  strict  safety  requirements,  where  violating  constraints  can  have  serious  consequences.  To  bridge  this  gap,  Safe  Reinforcement  Learning  (Safe  RL)  has  emerged  as  a  promising  approach  aimed  at  enabling  high-quality  decision-making  while  rigorously  enforcing  predefined  safety  constraints.  Despite  its  potential,  Safe  RL  not  only  inherits  the  fundamental  difficulties  of  standard  RL-such  as  high  sample  complexity  and  exploration-exploitation  trade-offs-but  also  introduces  additional  complexities.  These  include  (1)  the  need  to  carefully  balance  performance  objectives  with  safety  requirements  and  (2)  the  challenge  of  ensuring  algorithmic  convergence  under  nonconcave  objectives  and  nonconvex  constraints.Motivated  by  these  challenges,  this  thesis  aims  to  advance  the  theoretical  foundations  and  algorithmic  tools  for  Safe  RL,  focusing  on  the  framework  of  Constrained  Markov  Decision  Processes  (CMDPs).  We  develop  new  analytical  insights  and  propose  a  set  of  methods  tailored  to  three  representative  CMDP  problems:  (1)  Under  entropy  regularization,  we  introduce  a  dual  optimization  framework  that  takes  advantage  of  the  smoothness  induced  by  entropy  regularization.  This  leads  to  improved  convergence  guarantees  for  both  the  optimality  gap  and  constraint  violations.  (2)  When  the  objective  and  constraints  take  more  general  forms  than  the  standard  CMDP,  we  show  that  despite  losing  cumulative  structure,  hidden  concavity  can  be  exploited  to  design  policy-based  algorithms  with  provable  global  convergence  guarantees.  (3)  In  the  presence  of  multiple  agents,  we  come  up  with  a  distributed  primal-dual  algorithm  utilizing  shadow  rewards  and  communication  truncation  strategies,  effectively  addressing  scalability  issues  while  ensuring  convergence  to  first-order  stationary  points.
■590    ▼aSchool  code:  0028.
■650  4▼aIndustrial  engineering
■653    ▼aLagrangian  duality
■653    ▼aSafe  Reinforcement  Learning
■653    ▼aReinforcement  learning
■653    ▼aDecision-making
■653    ▼aAutonomous  driving
■690    ▼a0796
■690    ▼a0800
■690    ▼a0546
■71020▼aUniversity  of  California,  Berkeley▼bIndustrial  Engineering  &  Operations  Research.
■7730  ▼tDissertations  Abstracts  International▼g87-01B.
■790    ▼a0028
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357087▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF16376 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.