서브메뉴
검색
Theoretical Advances in Reinforcement Learning: Online Average-Reward and Offline Constrained Settings
Theoretical Advances in Reinforcement Learning: Online Average-Reward and Offline Constrained Settings
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105244
- ISBN
- 9798291569573
- DDC
- 310
- 저자명
- Hong, Kihyuk.
- 서명/저자
- Theoretical Advances in Reinforcement Learning: Online Average-Reward and Offline Constrained Settings
- 발행사항
- [Sl] : University of Michigan, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 148 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-03, Section: B.
- 주기사항
- Advisor: Tewari, Ambuj.
- 학위논문주기
- Thesis (Ph.D.)--University of Michigan, 2025.
- 초록/해제
- 요약This dissertation presents theoretical advancements in reinforcement learning (RL), focusing on two key settings: online RL under the infinite-horizon average-reward criterion and offline constrained RL under partial data coverage.The first part addresses the online setting, where the objective is to optimize long-run average rewards through interaction with the environment. A family of value-iteration-based algorithms is proposed by approximating the average-reward objective using a carefully tuned discounted surrogate. This part resolves an open problem by establishing a computationally efficient algorithm for linear Markov decision processes under a weak structural assumption. The proposed algorithm employs span-constrained value clipping and a decoupled planning strategy that mitigates statistical inefficiencies arising from the complexity of the function class.The second part is motivated by safety-critical applications and studies the offline setting, where the agent must learn a policy from a fixed dataset without further interaction. Primal-dual algorithms are developed for both linear MDPs and general function-approximation regimes, based on the linear programming formulation of RL. These methods are oracle-efficient and provably sample-efficient under partial data coverage. Moreover, they extend to the constrained RL setting, where the policy must satisfy additional safety constraints defined by auxiliary reward signals and threshold levels.Together, the contributions of this dissertation advance the theoretical foundations of reinforcement learning in settings that prioritize safety and long-term performance.
- 일반주제명
- Statistics
- 일반주제명
- Computer science
- 키워드
- Function class
- 기타저자
- University of Michigan Statistics
- 기본자료저록
- Dissertations Abstracts International. 87-03B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017359978
■00520260202105244
■006m o d
■007cr#unu||||||||
■020 ▼a9798291569573
■035 ▼a(MiAaPQ)AAI32272037
■035 ▼a(MiAaPQ)umichrackham006363
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a310
■1001 ▼aHong, Kihyuk.
■24510▼aTheoretical Advances in Reinforcement Learning: Online Average-Reward and Offline Constrained Settings
■260 ▼a[Sl]▼bUniversity of Michigan▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a148 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-03, Section: B.
■500 ▼aAdvisor: Tewari, Ambuj.
■5021 ▼aThesis (Ph.D.)--University of Michigan, 2025.
■520 ▼aThis dissertation presents theoretical advancements in reinforcement learning (RL), focusing on two key settings: online RL under the infinite-horizon average-reward criterion and offline constrained RL under partial data coverage.The first part addresses the online setting, where the objective is to optimize long-run average rewards through interaction with the environment. A family of value-iteration-based algorithms is proposed by approximating the average-reward objective using a carefully tuned discounted surrogate. This part resolves an open problem by establishing a computationally efficient algorithm for linear Markov decision processes under a weak structural assumption. The proposed algorithm employs span-constrained value clipping and a decoupled planning strategy that mitigates statistical inefficiencies arising from the complexity of the function class.The second part is motivated by safety-critical applications and studies the offline setting, where the agent must learn a policy from a fixed dataset without further interaction. Primal-dual algorithms are developed for both linear MDPs and general function-approximation regimes, based on the linear programming formulation of RL. These methods are oracle-efficient and provably sample-efficient under partial data coverage. Moreover, they extend to the constrained RL setting, where the policy must satisfy additional safety constraints defined by auxiliary reward signals and threshold levels.Together, the contributions of this dissertation advance the theoretical foundations of reinforcement learning in settings that prioritize safety and long-term performance.
■590 ▼aSchool code: 0127.
■650 4▼aStatistics
■650 4▼aComputer science
■653 ▼aReinforcement learning
■653 ▼aSafety constraints
■653 ▼aMarkov decision processes
■653 ▼aFunction class
■653 ▼aPrimal-dual algorithms
■690 ▼a0463
■690 ▼a0800
■690 ▼a0984
■71020▼aUniversity of Michigan▼bStatistics.
■7730 ▼tDissertations Abstracts International▼g87-03B.
■790 ▼a0127
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17359978▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


