본문

서브메뉴

Theoretical Advances in Reinforcement Learning: Online Average-Reward and Offline Constrained Settings
Theoretical Advances in Reinforcement Learning: Online Average-Reward and Offline Constrai...
Theoretical Advances in Reinforcement Learning: Online Average-Reward and Offline Constrained Settings

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202105244
ISBN  
9798291569573
DDC  
310
저자명  
Hong, Kihyuk.
서명/저자  
Theoretical Advances in Reinforcement Learning: Online Average-Reward and Offline Constrained Settings
발행사항  
[Sl] : University of Michigan, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
148 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-03, Section: B.
주기사항  
Advisor: Tewari, Ambuj.
학위논문주기  
Thesis (Ph.D.)--University of Michigan, 2025.
초록/해제  
요약This dissertation presents theoretical advancements in reinforcement learning (RL), focusing on two key settings: online RL under the infinite-horizon average-reward criterion and offline constrained RL under partial data coverage.The first part addresses the online setting, where the objective is to optimize long-run average rewards through interaction with the environment. A family of value-iteration-based algorithms is proposed by approximating the average-reward objective using a carefully tuned discounted surrogate. This part resolves an open problem by establishing a computationally efficient algorithm for linear Markov decision processes under a weak structural assumption. The proposed algorithm employs span-constrained value clipping and a decoupled planning strategy that mitigates statistical inefficiencies arising from the complexity of the function class.The second part is motivated by safety-critical applications and studies the offline setting, where the agent must learn a policy from a fixed dataset without further interaction. Primal-dual algorithms are developed for both linear MDPs and general function-approximation regimes, based on the linear programming formulation of RL. These methods are oracle-efficient and provably sample-efficient under partial data coverage. Moreover, they extend to the constrained RL setting, where the policy must satisfy additional safety constraints defined by auxiliary reward signals and threshold levels.Together, the contributions of this dissertation advance the theoretical foundations of reinforcement learning in settings that prioritize safety and long-term performance.
일반주제명  
Statistics
일반주제명  
Computer science
키워드  
Reinforcement learning
키워드  
Safety constraints
키워드  
Markov decision processes
키워드  
Function class
키워드  
Primal-dual algorithms
기타저자  
University of Michigan Statistics
기본자료저록  
Dissertations Abstracts International. 87-03B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017359978
■00520260202105244
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798291569573
■035    ▼a(MiAaPQ)AAI32272037
■035    ▼a(MiAaPQ)umichrackham006363
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a310
■1001  ▼aHong,  Kihyuk.
■24510▼aTheoretical  Advances  in  Reinforcement  Learning:  Online  Average-Reward  and  Offline  Constrained  Settings
■260    ▼a[Sl]▼bUniversity  of  Michigan▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a148  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-03,  Section:  B.
■500    ▼aAdvisor:  Tewari,  Ambuj.
■5021  ▼aThesis  (Ph.D.)--University  of  Michigan,  2025.
■520    ▼aThis  dissertation  presents  theoretical  advancements  in  reinforcement  learning  (RL),  focusing  on  two  key  settings:  online  RL  under  the  infinite-horizon  average-reward  criterion  and  offline  constrained  RL  under  partial  data  coverage.The  first  part  addresses  the  online  setting,  where  the  objective  is  to  optimize  long-run  average  rewards  through  interaction  with  the  environment.  A  family  of  value-iteration-based  algorithms  is  proposed  by  approximating  the  average-reward  objective  using  a  carefully  tuned  discounted  surrogate.  This  part  resolves  an  open  problem  by  establishing  a  computationally  efficient  algorithm  for  linear  Markov  decision  processes  under  a  weak  structural  assumption.  The  proposed  algorithm  employs  span-constrained  value  clipping  and  a  decoupled  planning  strategy  that  mitigates  statistical  inefficiencies  arising  from  the  complexity  of  the  function  class.The  second  part  is  motivated  by  safety-critical  applications  and  studies  the  offline  setting,  where  the  agent  must  learn  a  policy  from  a  fixed  dataset  without  further  interaction.  Primal-dual  algorithms  are  developed  for  both  linear  MDPs  and  general  function-approximation  regimes,  based  on  the  linear  programming  formulation  of  RL.  These  methods  are  oracle-efficient  and  provably  sample-efficient  under  partial  data  coverage.  Moreover,  they  extend  to  the  constrained  RL  setting,  where  the  policy  must  satisfy  additional  safety  constraints  defined  by  auxiliary  reward  signals  and  threshold  levels.Together,  the  contributions  of  this  dissertation  advance  the  theoretical  foundations  of  reinforcement  learning  in  settings  that  prioritize  safety  and  long-term  performance.
■590    ▼aSchool  code:  0127.
■650  4▼aStatistics
■650  4▼aComputer  science
■653    ▼aReinforcement  learning
■653    ▼aSafety  constraints
■653    ▼aMarkov  decision  processes
■653    ▼aFunction  class
■653    ▼aPrimal-dual  algorithms
■690    ▼a0463
■690    ▼a0800
■690    ▼a0984
■71020▼aUniversity  of  Michigan▼bStatistics.
■7730  ▼tDissertations  Abstracts  International▼g87-03B.
■790    ▼a0127
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17359978▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF18357 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.