본문

서브메뉴

The Role of Lookahead in Reinforcement Learning Algorithms
The Role of Lookahead in Reinforcement Learning Algorithms
The Role of Lookahead in Reinforcement Learning Algorithms

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202105701
ISBN  
9798263307868
DDC  
004
저자명  
Winnicki, Anna.
서명/저자  
The Role of Lookahead in Reinforcement Learning Algorithms
발행사항  
[Sl] : University of Illinois at Urbana-Champaign, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
106 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-05, Section: B.
주기사항  
Advisor: Srikant, R.
학위논문주기  
Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2024.
초록/해제  
요약State of the art reinforcement learning (RL) algorithms such as AlphaZero use lookahead, which is typically implemented using Monte Carlo Tree Search (MCTS). As the name suggests, lookahead simply means looking ahead several steps when computing the policy to be used. The fact that an H-step lookahead provides an O(αH), where α is the discount factor, approximate solution to the optimal policy is a somewhat trivial and well-known statement. What we have shown is a much stronger result: we have shown that lookahead leads to convergent learning algorithms while the same algorithms may diverge in the absence of lookahead. We have demonstrated these results for three different classes of RL algorithms: modified policy iteration with linear value function approximation, Monte Carlo with exploring starts, and policy iteration for zero-sum Markov games. We have also shown that lookahead can be efficiently implemented in the widely studied class of linear MDPs.
일반주제명  
Computer science
일반주제명  
Electrical engineering
일반주제명  
Computer engineering
키워드  
Reinforcement learning
키워드  
Markov decision processes
기타저자  
University of Illinois at Urbana-Champaign Electrical & Computer Eng
기본자료저록  
Dissertations Abstracts International. 87-05B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2024        us                              c    eng  d
■001000017361073
■00520260202105701
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798263307868
■035    ▼a(MiAaPQ)AAI32409879
■035    ▼a(MiAaPQ)124419
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aWinnicki,  Anna.
■24510▼aThe  Role  of  Lookahead  in  Reinforcement  Learning  Algorithms
■260    ▼a[Sl]▼bUniversity  of  Illinois  at  Urbana-Champaign▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a106  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-05,  Section:  B.
■500    ▼aAdvisor:  Srikant,  R.
■5021  ▼aThesis  (Ph.D.)--University  of  Illinois  at  Urbana-Champaign,  2024.
■520    ▼aState  of  the  art  reinforcement  learning  (RL)  algorithms  such  as  AlphaZero  use  lookahead,  which  is  typically  implemented  using  Monte  Carlo  Tree  Search  (MCTS).  As  the  name  suggests,  lookahead  simply  means  looking  ahead  several  steps  when  computing  the  policy  to  be  used.  The  fact  that  an  H-step  lookahead  provides  an  O(αH),  where  α  is  the  discount  factor,  approximate  solution  to  the  optimal  policy  is  a  somewhat  trivial  and  well-known  statement.  What  we  have  shown  is  a  much  stronger  result:  we  have  shown  that  lookahead  leads  to  convergent  learning  algorithms  while  the  same  algorithms  may  diverge  in  the  absence  of  lookahead.  We  have  demonstrated  these  results  for  three  different  classes  of  RL  algorithms:  modified  policy  iteration  with  linear  value  function  approximation,  Monte  Carlo  with  exploring  starts,  and  policy  iteration  for  zero-sum  Markov  games.  We  have  also  shown  that  lookahead  can  be  efficiently  implemented  in  the  widely  studied  class  of  linear  MDPs.
■590    ▼aSchool  code:  0090.
■650  4▼aComputer  science
■650  4▼aElectrical  engineering
■650  4▼aComputer  engineering
■653    ▼aReinforcement  learning
■653    ▼aMarkov  decision  processes
■690    ▼a0544
■690    ▼a0984
■690    ▼a0464
■71020▼aUniversity  of  Illinois  at  Urbana-Champaign▼bElectrical  &  Computer  Eng.
■7730  ▼tDissertations  Abstracts  International▼g87-05B.
■790    ▼a0090
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17361073▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF16750 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.