서브메뉴
검색
The Role of Lookahead in Reinforcement Learning Algorithms
The Role of Lookahead in Reinforcement Learning Algorithms
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105701
- ISBN
- 9798263307868
- DDC
- 004
- 저자명
- Winnicki, Anna.
- 서명/저자
- The Role of Lookahead in Reinforcement Learning Algorithms
- 발행사항
- [Sl] : University of Illinois at Urbana-Champaign, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 106 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-05, Section: B.
- 주기사항
- Advisor: Srikant, R.
- 학위논문주기
- Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2024.
- 초록/해제
- 요약State of the art reinforcement learning (RL) algorithms such as AlphaZero use lookahead, which is typically implemented using Monte Carlo Tree Search (MCTS). As the name suggests, lookahead simply means looking ahead several steps when computing the policy to be used. The fact that an H-step lookahead provides an O(αH), where α is the discount factor, approximate solution to the optimal policy is a somewhat trivial and well-known statement. What we have shown is a much stronger result: we have shown that lookahead leads to convergent learning algorithms while the same algorithms may diverge in the absence of lookahead. We have demonstrated these results for three different classes of RL algorithms: modified policy iteration with linear value function approximation, Monte Carlo with exploring starts, and policy iteration for zero-sum Markov games. We have also shown that lookahead can be efficiently implemented in the widely studied class of linear MDPs.
- 일반주제명
- Computer science
- 일반주제명
- Electrical engineering
- 일반주제명
- Computer engineering
- 기타저자
- University of Illinois at Urbana-Champaign Electrical & Computer Eng
- 기본자료저록
- Dissertations Abstracts International. 87-05B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2024 us c eng d■001000017361073
■00520260202105701
■006m o d
■007cr#unu||||||||
■020 ▼a9798263307868
■035 ▼a(MiAaPQ)AAI32409879
■035 ▼a(MiAaPQ)124419
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aWinnicki, Anna.
■24510▼aThe Role of Lookahead in Reinforcement Learning Algorithms
■260 ▼a[Sl]▼bUniversity of Illinois at Urbana-Champaign▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a106 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-05, Section: B.
■500 ▼aAdvisor: Srikant, R.
■5021 ▼aThesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2024.
■520 ▼aState of the art reinforcement learning (RL) algorithms such as AlphaZero use lookahead, which is typically implemented using Monte Carlo Tree Search (MCTS). As the name suggests, lookahead simply means looking ahead several steps when computing the policy to be used. The fact that an H-step lookahead provides an O(αH), where α is the discount factor, approximate solution to the optimal policy is a somewhat trivial and well-known statement. What we have shown is a much stronger result: we have shown that lookahead leads to convergent learning algorithms while the same algorithms may diverge in the absence of lookahead. We have demonstrated these results for three different classes of RL algorithms: modified policy iteration with linear value function approximation, Monte Carlo with exploring starts, and policy iteration for zero-sum Markov games. We have also shown that lookahead can be efficiently implemented in the widely studied class of linear MDPs.
■590 ▼aSchool code: 0090.
■650 4▼aComputer science
■650 4▼aElectrical engineering
■650 4▼aComputer engineering
■653 ▼aReinforcement learning
■653 ▼aMarkov decision processes
■690 ▼a0544
■690 ▼a0984
■690 ▼a0464
■71020▼aUniversity of Illinois at Urbana-Champaign▼bElectrical & Computer Eng.
■7730 ▼tDissertations Abstracts International▼g87-05B.
■790 ▼a0090
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17361073▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


