본문

서브메뉴

The Role of Lookahead in Reinforcement Learning Algorithms
The Role of Lookahead in Reinforcement Learning Algorithms
The Role of Lookahead in Reinforcement Learning Algorithms

Detailed Information

자료유형  
 학위논문 서양
최종처리일시  
20260202105701
ISBN  
9798263307868
DDC  
004
저자명  
Winnicki, Anna.
서명/저자  
The Role of Lookahead in Reinforcement Learning Algorithms
발행사항  
[Sl] : University of Illinois at Urbana-Champaign, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
106 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-05, Section: B.
주기사항  
Advisor: Srikant, R.
학위논문주기  
Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2024.
초록/해제  
요약State of the art reinforcement learning (RL) algorithms such as AlphaZero use lookahead, which is typically implemented using Monte Carlo Tree Search (MCTS). As the name suggests, lookahead simply means looking ahead several steps when computing the policy to be used. The fact that an H-step lookahead provides an O(αH), where α is the discount factor, approximate solution to the optimal policy is a somewhat trivial and well-known statement. What we have shown is a much stronger result: we have shown that lookahead leads to convergent learning algorithms while the same algorithms may diverge in the absence of lookahead. We have demonstrated these results for three different classes of RL algorithms: modified policy iteration with linear value function approximation, Monte Carlo with exploring starts, and policy iteration for zero-sum Markov games. We have also shown that lookahead can be efficiently implemented in the widely studied class of linear MDPs.
일반주제명  
Computer science
일반주제명  
Electrical engineering
일반주제명  
Computer engineering
키워드  
Reinforcement learning
키워드  
Markov decision processes
기타저자  
University of Illinois at Urbana-Champaign Electrical & Computer Eng
기본자료저록  
Dissertations Abstracts International. 87-05B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2024        us                              c    eng  d
■001000017361073
■00520260202105701
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798263307868
■035    ▼a(MiAaPQ)AAI32409879
■035    ▼a(MiAaPQ)124419
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aWinnicki,  Anna.
■24510▼aThe  Role  of  Lookahead  in  Reinforcement  Learning  Algorithms
■260    ▼a[Sl]▼bUniversity  of  Illinois  at  Urbana-Champaign▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a106  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-05,  Section:  B.
■500    ▼aAdvisor:  Srikant,  R.
■5021  ▼aThesis  (Ph.D.)--University  of  Illinois  at  Urbana-Champaign,  2024.
■520    ▼aState  of  the  art  reinforcement  learning  (RL)  algorithms  such  as  AlphaZero  use  lookahead,  which  is  typically  implemented  using  Monte  Carlo  Tree  Search  (MCTS).  As  the  name  suggests,  lookahead  simply  means  looking  ahead  several  steps  when  computing  the  policy  to  be  used.  The  fact  that  an  H-step  lookahead  provides  an  O(αH),  where  α  is  the  discount  factor,  approximate  solution  to  the  optimal  policy  is  a  somewhat  trivial  and  well-known  statement.  What  we  have  shown  is  a  much  stronger  result:  we  have  shown  that  lookahead  leads  to  convergent  learning  algorithms  while  the  same  algorithms  may  diverge  in  the  absence  of  lookahead.  We  have  demonstrated  these  results  for  three  different  classes  of  RL  algorithms:  modified  policy  iteration  with  linear  value  function  approximation,  Monte  Carlo  with  exploring  starts,  and  policy  iteration  for  zero-sum  Markov  games.  We  have  also  shown  that  lookahead  can  be  efficiently  implemented  in  the  widely  studied  class  of  linear  MDPs.
■590    ▼aSchool  code:  0090.
■650  4▼aComputer  science
■650  4▼aElectrical  engineering
■650  4▼aComputer  engineering
■653    ▼aReinforcement  learning
■653    ▼aMarkov  decision  processes
■690    ▼a0544
■690    ▼a0984
■690    ▼a0464
■71020▼aUniversity  of  Illinois  at  Urbana-Champaign▼bElectrical  &  Computer  Eng.
■7730  ▼tDissertations  Abstracts  International▼g87-05B.
■790    ▼a0090
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17361073▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

Preview

Export

ChatGPT Discussion

AI Recommended Related Books


    New Books MORE
    Statistics for the past 3 years. Go to brief

    Info Détail de la recherche.

    • Réservation
    • n'existe pas
    • My Folder
    • Demande Première utilisation
    • Non-Book Loan Application
    • Nighttime Book Loan Application
    Matériel
    Reg No. Call No. emplacement Status Lend Info
    TF16750 전자도서 대출가능 My Folder 부재도서신고 비도서대출신청 야간 도서대출신청

    * Les réservations sont disponibles dans le livre d'emprunt. Pour faire des réservations, S'il vous plaît cliquer sur le bouton de réservation

    Books borrowed together with this book

    Related Popular Books

    Available after logging in.