서브메뉴
검색
Deciding What to Learn in Complex Environments
Deciding What to Learn in Complex Environments
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211152744
- ISBN
- 9798342108409
- DDC
- 658
- 서명/저자
- Deciding What to Learn in Complex Environments
- 발행사항
- [Sl] : Stanford University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 118 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-04, Section: B.
- 주기사항
- Advisor: Finn, Chelsea;Roy, Benjamin Van.
- 학위논문주기
- Thesis (Ph.D.)--Stanford University, 2024.
- 초록/해제
- 요약Reinforcement learning is the paradigm of machine learning dedicated to sequential decision-making problems. Like many other areas of machine learning and statistics, there is often a prevalent concern over data efficiency; that is, how much trial-and-error interaction data does a sequential decision-making agent require in order to learn desired behaviors? One of the key obstacles to data-efficient reinforcement learning is the challenge of exploration, whereby a sequential decision-making agent must balance between gaining new knowledge about an environment and exploiting current knowledge to maximize near-term performance. The traditional literature on balancing exploration and exploitation focuses on environments in which an agent can approach optimal performance within a relevant time frame. However, modern artificial decision-making agents engage with complex environments, such as the World Wide Web, in which there is no hope of approaching optimal performance within any relevant time frame.This dissertation focuses on developing principled, practical approaches for addressing the exploration problem in complex environments. Our methods focus on the simple observation that, rather than endeavor to obtain enough information for achieving optimal behavior, an agent confronted with such a complex environment should instead target a modest corpus of information that, while capable of facilitating behavioral improvement, is itself insufficient to enable near-optimal performance. We design an agent that modulates exploration in this way and provide both a theoretical as well as an empirical analysis of its behavior. In effect, at each time period, this agent decides what to learn so as to strike a desired trade-off between information requirements and performance. As we elucidate in this dissertation, central to the design of such an agent are classic tools from information theory and lossy compression, which not only facilitate principled theoretical guarantees but also remain amenable to practical implementation at scale.
- 일반주제명
- Behavior
- 일반주제명
- Deep learning
- 일반주제명
- Cognitive science
- 일반주제명
- Decision making
- 일반주제명
- Design
- 일반주제명
- Information processing
- 일반주제명
- Codes
- 일반주제명
- Core curriculum
- 일반주제명
- Information theory
- 일반주제명
- Entropy
- 일반주제명
- Applied mathematics
- 일반주제명
- Curriculum development
- 기타저자
- Stanford University.
- 기본자료저록
- Dissertations Abstracts International. 86-04B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017163716
■00520250211152744
■006m o d
■007cr#unu||||||||
■020 ▼a9798342108409
■035 ▼a(MiAaPQ)AAI31520266
■035 ▼a(MiAaPQ)Stanfordfh741vv5133
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a658
■1001 ▼aArumugam, Dilip Srimal.
■24510▼aDeciding What to Learn in Complex Environments
■260 ▼a[Sl]▼bStanford University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a118 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-04, Section: B.
■500 ▼aAdvisor: Finn, Chelsea;Roy, Benjamin Van.
■5021 ▼aThesis (Ph.D.)--Stanford University, 2024.
■520 ▼aReinforcement learning is the paradigm of machine learning dedicated to sequential decision-making problems. Like many other areas of machine learning and statistics, there is often a prevalent concern over data efficiency; that is, how much trial-and-error interaction data does a sequential decision-making agent require in order to learn desired behaviors? One of the key obstacles to data-efficient reinforcement learning is the challenge of exploration, whereby a sequential decision-making agent must balance between gaining new knowledge about an environment and exploiting current knowledge to maximize near-term performance. The traditional literature on balancing exploration and exploitation focuses on environments in which an agent can approach optimal performance within a relevant time frame. However, modern artificial decision-making agents engage with complex environments, such as the World Wide Web, in which there is no hope of approaching optimal performance within any relevant time frame.This dissertation focuses on developing principled, practical approaches for addressing the exploration problem in complex environments. Our methods focus on the simple observation that, rather than endeavor to obtain enough information for achieving optimal behavior, an agent confronted with such a complex environment should instead target a modest corpus of information that, while capable of facilitating behavioral improvement, is itself insufficient to enable near-optimal performance. We design an agent that modulates exploration in this way and provide both a theoretical as well as an empirical analysis of its behavior. In effect, at each time period, this agent decides what to learn so as to strike a desired trade-off between information requirements and performance. As we elucidate in this dissertation, central to the design of such an agent are classic tools from information theory and lossy compression, which not only facilitate principled theoretical guarantees but also remain amenable to practical implementation at scale.
■590 ▼aSchool code: 0212.
■650 4▼aBehavior
■650 4▼aDeep learning
■650 4▼aCognitive science
■650 4▼aDecision making
■650 4▼aDesign
■650 4▼aInformation processing
■650 4▼aCodes
■650 4▼aCore curriculum
■650 4▼aInformation theory
■650 4▼aEntropy
■650 4▼aApplied mathematics
■650 4▼aCurriculum development
■690 ▼a0389
■690 ▼a0800
■690 ▼a0364
■690 ▼a0727
■71020▼aStanford University.
■7730 ▼tDissertations Abstracts International▼g86-04B.
■790 ▼a0212
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17163716▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


