본문

서브메뉴

Reinforcement Learning Beyond Rewards: Decision-Making in the Language of Visitation Distributions
Reinforcement Learning Beyond Rewards: Decision-Making in the Language of Visitation Distr...
Reinforcement Learning Beyond Rewards: Decision-Making in the Language of Visitation Distributions

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260311091515.5
ISBN  
9798270229436
DDC  
006.31
저자명  
Sikchi, Harshit Sushil
서명/저자  
Reinforcement Learning Beyond Rewards: Decision-Making in the Language of Visitation Distributions / Harshit Sushil Sikchi
발행사항  
[Sl] : The University of Texas at Austin, 2025
형태사항  
1 electronic resource (423 pages)
주기사항  
Source: Dissertations Abstracts International, Volume: 87-06, Section: B.
주기사항  
Advisors: Zhang, Amy; Niekum, Scott David Committee members: Bellemare, Marc G.; Stone, Peter; Zhu, Yuke.
학위논문주기  
- Ph.D. : The University of Texas at Austin, 2025.
초록/해제  
요약Reinforcement Learning (RL) is traditionally framed as the problem of finding a policy that maximizes the cumulative reward in the environment. The generality of the RL framework rests on the reward function being a universal way to specify a decision-making task to an agent. While this notion of the universality of reward function has been debated, reward functions can often be inconvenient for task specification. Small changes in reward function can completely change optimal policy and it has been evidenced that humans frequently make mistakes when specifying tasks as rewards, resulting in a policy that is misaligned with the human intention. Alternatively, the environment is abundant with different forms of learning signals, and this thesis aims to investigate an alternate framework for decision-making that allows for a unified way to directly learn from a variety of learning signals not limited to reward functions.The core idea proposed in this thesis is to focus on visitation distributions as the central object of optimization- the future state-action distribution of any policy when interacting with the environment. Specifically, we show the generality of this framework by providing a unified set of algorithms that are able to learn from the following learning signals - rewards, goals, expert demonstrations, and action-free demonstrations. Our algorithms simplify optimization and forgo reward inference when learning from other signals and directly attempt to learn optimal policies.While environmental signals can greatly influence learning a particular task, a bulk of an interactive agent's experience in the environment may not have any associated learning signal. Even for this case, we hypothesize that the future state-action visitation distribution of an agent captures information necessary for decision making. This insight allows us to propose a self-supervised objective for decision-making that learns representations by learning to represent all possible visitations in the environment using offline datasets without any learning signals. We show that such an unsupervised learning approach can give rise to general-purpose RL agents that can perform any task specified by a reward function, video demonstration, or language instruction near-optimally without any test-time planning or learning. Finally, the thesis concludes by providing a solution to quickly adapt these near-optimal policies given by unsupervised RL agents rapidly for a test-time reward function.
언어주기  
English
일반주제명  
Computer science
키워드  
Reinforcement Learning
키워드  
Visitation distributions
키워드  
Self-supervised objective
기타저자  
The University of Texas at Austin Computer Science
기본자료저록  
Dissertations Abstracts International. 87-06B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260311s2025        us                                    eng  d
■001000017361146
■00520260311091515.5
■006m          o    d                
■007cr|nu||||||||
■020    ▼a9798270229436
■040    ▼aMiAaPQD▼beng▼cMiAaPQD▼erda
■082    ▼a006.31
■1001  ▼aSikchi,  Harshit  Sushil▼eauthor.
■24510▼aReinforcement  Learning  Beyond  Rewards:  Decision-Making  in  the  Language  of  Visitation  Distributions  ▼cHarshit  Sushil  Sikchi
■260    ▼a[Sl]▼bThe  University  of  Texas  at  Austin▼c2025
■264  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a1  electronic  resource  (423  pages)
■336    ▼atext▼btxt▼2rdacontent
■337    ▼acomputer▼bc▼2rdamedia
■338    ▼aonline  resource▼bcr▼2rdacarrier
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-06,  Section:  B.
■500    ▼aAdvisors:  Zhang,  Amy;  Niekum,  Scott  David    Committee  members:  Bellemare,  Marc  G.;  Stone,  Peter;  Zhu,  Yuke.
■5021  ▼bPh.D.▼cThe  University  of  Texas  at  Austin▼d2025.
■520    ▼aReinforcement  Learning  (RL)  is  traditionally  framed  as  the  problem  of  finding  a  policy  that  maximizes  the  cumulative  reward  in  the  environment.  The  generality  of  the  RL  framework  rests  on  the  reward  function  being  a  universal  way  to  specify  a  decision-making  task  to  an  agent.  While  this  notion  of  the  universality  of  reward  function  has  been  debated,  reward  functions  can  often  be  inconvenient  for  task  specification.  Small  changes  in  reward  function  can  completely  change  optimal  policy  and  it  has  been  evidenced  that  humans  frequently  make  mistakes  when  specifying  tasks  as  rewards,  resulting  in  a  policy  that  is  misaligned  with  the  human  intention.  Alternatively,  the  environment  is  abundant  with  different  forms  of  learning  signals,  and  this  thesis  aims  to  investigate  an  alternate  framework  for  decision-making  that  allows  for  a  unified  way  to  directly  learn  from  a  variety  of  learning  signals  not  limited  to  reward  functions.The  core  idea  proposed  in  this  thesis  is  to  focus  on  visitation  distributions  as  the  central  object  of  optimization-  the  future  state-action  distribution  of  any  policy  when  interacting  with  the  environment.  Specifically,  we  show  the  generality  of  this  framework  by  providing  a  unified  set  of  algorithms  that  are  able  to  learn  from  the  following  learning  signals  -  rewards,  goals,  expert  demonstrations,  and  action-free  demonstrations.  Our  algorithms  simplify  optimization  and  forgo  reward  inference  when  learning  from  other  signals  and  directly  attempt  to  learn  optimal  policies.While  environmental  signals  can  greatly  influence  learning  a  particular  task,  a  bulk  of  an  interactive  agent's  experience  in  the  environment  may  not  have  any  associated  learning  signal.  Even  for  this  case,  we  hypothesize  that  the  future  state-action  visitation  distribution  of  an  agent  captures  information  necessary  for  decision  making.  This  insight  allows  us  to  propose  a  self-supervised  objective  for  decision-making  that  learns  representations  by  learning  to  represent  all  possible  visitations  in  the  environment  using  offline  datasets  without  any  learning  signals.  We  show  that  such  an  unsupervised  learning  approach  can  give  rise  to  general-purpose  RL  agents  that  can  perform  any  task  specified  by  a  reward  function,  video  demonstration,  or  language  instruction  near-optimally  without  any  test-time  planning  or  learning.  Finally,  the  thesis  concludes  by  providing  a  solution  to  quickly  adapt  these  near-optimal  policies  given  by  unsupervised  RL  agents  rapidly  for  a  test-time  reward  function.
■546    ▼aEnglish
■590    ▼aSchool  code:  0227
■650  4▼aComputer  science
■653    ▼aReinforcement  Learning
■653    ▼aVisitation  distributions
■653    ▼aSelf-supervised  objective
■7102  ▼aThe  University  of  Texas  at  Austin▼bComputer  Science.▼edegree  granting  institution.
■7201  ▼aZhang,  Amy▼edegree  supervisor.
■7201  ▼aNiekum,  Scott  David▼edegree  supervisor.
■7730  ▼tDissertations  Abstracts  International▼g87-06B.
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17361146▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF18232 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.