본문

서브메뉴

Deciding What to Learn in Complex Environments
Deciding What to Learn in Complex Environments
Deciding What to Learn in Complex Environments

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211152744
ISBN  
9798342108409
DDC  
658
저자명  
Arumugam, Dilip Srimal.
서명/저자  
Deciding What to Learn in Complex Environments
발행사항  
[Sl] : Stanford University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
118 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-04, Section: B.
주기사항  
Advisor: Finn, Chelsea;Roy, Benjamin Van.
학위논문주기  
Thesis (Ph.D.)--Stanford University, 2024.
초록/해제  
요약Reinforcement learning is the paradigm of machine learning dedicated to sequential decision-making problems. Like many other areas of machine learning and statistics, there is often a prevalent concern over data efficiency; that is, how much trial-and-error interaction data does a sequential decision-making agent require in order to learn desired behaviors? One of the key obstacles to data-efficient reinforcement learning is the challenge of exploration, whereby a sequential decision-making agent must balance between gaining new knowledge about an environment and exploiting current knowledge to maximize near-term performance. The traditional literature on balancing exploration and exploitation focuses on environments in which an agent can approach optimal performance within a relevant time frame. However, modern artificial decision-making agents engage with complex environments, such as the World Wide Web, in which there is no hope of approaching optimal performance within any relevant time frame.This dissertation focuses on developing principled, practical approaches for addressing the exploration problem in complex environments. Our methods focus on the simple observation that, rather than endeavor to obtain enough information for achieving optimal behavior, an agent confronted with such a complex environment should instead target a modest corpus of information that, while capable of facilitating behavioral improvement, is itself insufficient to enable near-optimal performance. We design an agent that modulates exploration in this way and provide both a theoretical as well as an empirical analysis of its behavior. In effect, at each time period, this agent decides what to learn so as to strike a desired trade-off between information requirements and performance. As we elucidate in this dissertation, central to the design of such an agent are classic tools from information theory and lossy compression, which not only facilitate principled theoretical guarantees but also remain amenable to practical implementation at scale.
일반주제명  
Behavior
일반주제명  
Deep learning
일반주제명  
Cognitive science
일반주제명  
Decision making
일반주제명  
Design
일반주제명  
Information processing
일반주제명  
Codes
일반주제명  
Core curriculum
일반주제명  
Information theory
일반주제명  
Entropy
일반주제명  
Applied mathematics
일반주제명  
Curriculum development
기타저자  
Stanford University.
기본자료저록  
Dissertations Abstracts International. 86-04B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017163716
■00520250211152744
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798342108409
■035    ▼a(MiAaPQ)AAI31520266
■035    ▼a(MiAaPQ)Stanfordfh741vv5133
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a658
■1001  ▼aArumugam,  Dilip  Srimal.
■24510▼aDeciding  What  to  Learn  in  Complex  Environments
■260    ▼a[Sl]▼bStanford  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a118  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-04,  Section:  B.
■500    ▼aAdvisor:  Finn,  Chelsea;Roy,  Benjamin  Van.
■5021  ▼aThesis  (Ph.D.)--Stanford  University,  2024.
■520    ▼aReinforcement  learning  is  the  paradigm  of  machine  learning  dedicated  to  sequential  decision-making  problems.  Like  many  other  areas  of  machine  learning  and  statistics,  there  is  often  a  prevalent  concern  over  data  efficiency;  that  is,  how  much  trial-and-error  interaction  data  does  a  sequential  decision-making  agent  require  in  order  to  learn  desired  behaviors?  One  of  the  key  obstacles  to  data-efficient  reinforcement  learning  is  the  challenge  of  exploration,  whereby  a  sequential  decision-making  agent  must  balance  between  gaining  new  knowledge  about  an  environment  and  exploiting  current  knowledge  to  maximize  near-term  performance.  The  traditional  literature  on  balancing  exploration  and  exploitation  focuses  on  environments  in  which  an  agent  can  approach  optimal  performance  within  a  relevant  time  frame.  However,  modern  artificial  decision-making  agents  engage  with  complex  environments,  such  as  the  World  Wide  Web,  in  which  there  is  no  hope  of  approaching  optimal  performance  within  any  relevant  time  frame.This  dissertation  focuses  on  developing  principled,  practical  approaches  for  addressing  the  exploration  problem  in  complex  environments.  Our  methods  focus  on  the  simple  observation  that,  rather  than  endeavor  to  obtain  enough  information  for  achieving  optimal  behavior,  an  agent  confronted  with  such  a  complex  environment  should  instead  target  a  modest  corpus  of  information  that,  while  capable  of  facilitating  behavioral  improvement,  is  itself  insufficient  to  enable  near-optimal  performance.  We  design  an  agent  that  modulates  exploration  in  this  way  and  provide  both  a  theoretical  as  well  as  an  empirical  analysis  of  its  behavior.  In  effect,  at  each  time  period,  this  agent  decides  what  to  learn  so  as  to  strike  a  desired  trade-off  between  information  requirements  and  performance.  As  we  elucidate  in  this  dissertation,  central  to  the  design  of  such  an  agent  are  classic  tools  from  information  theory  and  lossy  compression,  which  not  only  facilitate  principled  theoretical  guarantees  but  also  remain  amenable  to  practical  implementation  at  scale.
■590    ▼aSchool  code:  0212.
■650  4▼aBehavior
■650  4▼aDeep  learning
■650  4▼aCognitive  science
■650  4▼aDecision  making
■650  4▼aDesign
■650  4▼aInformation  processing
■650  4▼aCodes
■650  4▼aCore  curriculum
■650  4▼aInformation  theory
■650  4▼aEntropy
■650  4▼aApplied  mathematics
■650  4▼aCurriculum  development
■690    ▼a0389
■690    ▼a0800
■690    ▼a0364
■690    ▼a0727
■71020▼aStanford  University.
■7730  ▼tDissertations  Abstracts  International▼g86-04B.
■790    ▼a0212
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17163716▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF10661 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.