서브메뉴
검색
Reinforcement Learning Beyond Rewards: Decision-Making in the Language of Visitation Distributions
Reinforcement Learning Beyond Rewards: Decision-Making in the Language of Visitation Distributions
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260311091515.5
- ISBN
- 9798270229436
- DDC
- 006.31
- 서명/저자
- Reinforcement Learning Beyond Rewards: Decision-Making in the Language of Visitation Distributions / Harshit Sushil Sikchi
- 발행사항
- [Sl] : The University of Texas at Austin, 2025
- 형태사항
- 1 electronic resource (423 pages)
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-06, Section: B.
- 주기사항
- Advisors: Zhang, Amy; Niekum, Scott David Committee members: Bellemare, Marc G.; Stone, Peter; Zhu, Yuke.
- 학위논문주기
- - Ph.D. : The University of Texas at Austin, 2025.
- 초록/해제
- 요약Reinforcement Learning (RL) is traditionally framed as the problem of finding a policy that maximizes the cumulative reward in the environment. The generality of the RL framework rests on the reward function being a universal way to specify a decision-making task to an agent. While this notion of the universality of reward function has been debated, reward functions can often be inconvenient for task specification. Small changes in reward function can completely change optimal policy and it has been evidenced that humans frequently make mistakes when specifying tasks as rewards, resulting in a policy that is misaligned with the human intention. Alternatively, the environment is abundant with different forms of learning signals, and this thesis aims to investigate an alternate framework for decision-making that allows for a unified way to directly learn from a variety of learning signals not limited to reward functions.The core idea proposed in this thesis is to focus on visitation distributions as the central object of optimization- the future state-action distribution of any policy when interacting with the environment. Specifically, we show the generality of this framework by providing a unified set of algorithms that are able to learn from the following learning signals - rewards, goals, expert demonstrations, and action-free demonstrations. Our algorithms simplify optimization and forgo reward inference when learning from other signals and directly attempt to learn optimal policies.While environmental signals can greatly influence learning a particular task, a bulk of an interactive agent's experience in the environment may not have any associated learning signal. Even for this case, we hypothesize that the future state-action visitation distribution of an agent captures information necessary for decision making. This insight allows us to propose a self-supervised objective for decision-making that learns representations by learning to represent all possible visitations in the environment using offline datasets without any learning signals. We show that such an unsupervised learning approach can give rise to general-purpose RL agents that can perform any task specified by a reward function, video demonstration, or language instruction near-optimally without any test-time planning or learning. Finally, the thesis concludes by providing a solution to quickly adapt these near-optimal policies given by unsupervised RL agents rapidly for a test-time reward function.
- 언어주기
- English
- 일반주제명
- Computer science
- 기타저자
- The University of Texas at Austin Computer Science
- 기본자료저록
- Dissertations Abstracts International. 87-06B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260311s2025 us eng d■001000017361146
■00520260311091515.5
■006m o d
■007cr|nu||||||||
■020 ▼a9798270229436
■040 ▼aMiAaPQD▼beng▼cMiAaPQD▼erda
■082 ▼a006.31
■1001 ▼aSikchi, Harshit Sushil▼eauthor.
■24510▼aReinforcement Learning Beyond Rewards: Decision-Making in the Language of Visitation Distributions ▼cHarshit Sushil Sikchi
■260 ▼a[Sl]▼bThe University of Texas at Austin▼c2025
■264 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a1 electronic resource (423 pages)
■336 ▼atext▼btxt▼2rdacontent
■337 ▼acomputer▼bc▼2rdamedia
■338 ▼aonline resource▼bcr▼2rdacarrier
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-06, Section: B.
■500 ▼aAdvisors: Zhang, Amy; Niekum, Scott David Committee members: Bellemare, Marc G.; Stone, Peter; Zhu, Yuke.
■5021 ▼bPh.D.▼cThe University of Texas at Austin▼d2025.
■520 ▼aReinforcement Learning (RL) is traditionally framed as the problem of finding a policy that maximizes the cumulative reward in the environment. The generality of the RL framework rests on the reward function being a universal way to specify a decision-making task to an agent. While this notion of the universality of reward function has been debated, reward functions can often be inconvenient for task specification. Small changes in reward function can completely change optimal policy and it has been evidenced that humans frequently make mistakes when specifying tasks as rewards, resulting in a policy that is misaligned with the human intention. Alternatively, the environment is abundant with different forms of learning signals, and this thesis aims to investigate an alternate framework for decision-making that allows for a unified way to directly learn from a variety of learning signals not limited to reward functions.The core idea proposed in this thesis is to focus on visitation distributions as the central object of optimization- the future state-action distribution of any policy when interacting with the environment. Specifically, we show the generality of this framework by providing a unified set of algorithms that are able to learn from the following learning signals - rewards, goals, expert demonstrations, and action-free demonstrations. Our algorithms simplify optimization and forgo reward inference when learning from other signals and directly attempt to learn optimal policies.While environmental signals can greatly influence learning a particular task, a bulk of an interactive agent's experience in the environment may not have any associated learning signal. Even for this case, we hypothesize that the future state-action visitation distribution of an agent captures information necessary for decision making. This insight allows us to propose a self-supervised objective for decision-making that learns representations by learning to represent all possible visitations in the environment using offline datasets without any learning signals. We show that such an unsupervised learning approach can give rise to general-purpose RL agents that can perform any task specified by a reward function, video demonstration, or language instruction near-optimally without any test-time planning or learning. Finally, the thesis concludes by providing a solution to quickly adapt these near-optimal policies given by unsupervised RL agents rapidly for a test-time reward function.
■546 ▼aEnglish
■590 ▼aSchool code: 0227
■650 4▼aComputer science
■653 ▼aReinforcement Learning
■653 ▼aVisitation distributions
■653 ▼aSelf-supervised objective
■7102 ▼aThe University of Texas at Austin▼bComputer Science.▼edegree granting institution.
■7201 ▼aZhang, Amy▼edegree supervisor.
■7201 ▼aNiekum, Scott David▼edegree supervisor.
■7730 ▼tDissertations Abstracts International▼g87-06B.
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17361146▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


