서브메뉴
검색
Environment Generation for Autonomous Agents for Sequential Decision Making
Environment Generation for Autonomous Agents for Sequential Decision Making
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211152753
- ISBN
- 9798384448891
- DDC
- 004
- 서명/저자
- Environment Generation for Autonomous Agents for Sequential Decision Making
- 발행사항
- [Sl] : University of California, Berkeley, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 115 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-04, Section: B.
- 주기사항
- Advisor: Stoica, Ion.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Berkeley, 2024.
- 초록/해제
- 요약Autonomous agents have seen tremendous advancements in solving sequential decision-making problems in recent years, primarily driven by Reinforcement Learning (RL), and more recently, by Large Generative Models. The capability of these autonomous agents depends crucially on the quality and diversity of the learning environments they are trained in. This thesis presents research on designing frameworks and algorithms to formulate and systematically generate environments that improve the generalization capabilities of autonomous agents in solving sequential decision-making tasks. First, we explore the benefits of human-guided programmatic environment generation for training, testing, and debugging autonomous agents in complex real-time strategic (RTS) environments. We present a novel framework that, for the first time, demonstrates the benefits of using scenario specification languages (e.g., SCENIC) for systematic modeling and generation of realistic and diverse RTS RL environments (e.g., Soccer). Next, we discuss a class of algorithms called adaptive teacher Unsupervised Environment Design (UED), which automatically generates training tasks with an RL teacher agent. UED shows promising zero-shot generalization by simultaneously learning a task distribution (i.e., curriculum) and agent policies on the generated tasks. This is a non-stationary process where the task distribution evolves along with agent policies, creating instability over time. While prior works demonstrated the potential of such approaches, training the teacher remained a practical challenge. To this end, we introduce Curriculum Learning via Unsupervised Task Representation Learning (CLUTR): a novel unsupervised curriculum learning algorithm that decouples task representation and curriculum learning into a two-stage optimization to solve the training instability by pretraining a latent task manifold. Following that, we present Multi-Modal Reasoning and Critique for web navigation (MMRC), which introduces augmented environments with multimodal critic agents to enhance the performance of Large Foundational Multimodal Language agents on autonomous web navigation tasks. Together, these approaches portray the importance and usefulness of environment formulation and generation encompassing traditional RL-based and contemporary LLM-based agents.
- 일반주제명
- Computer science
- 일반주제명
- Computer engineering
- 기타저자
- University of California, Berkeley Computer Science
- 기본자료저록
- Dissertations Abstracts International. 86-04B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017163786
■00520250211152753
■006m o d
■007cr#unu||||||||
■020 ▼a9798384448891
■035 ▼a(MiAaPQ)AAI31555572
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aAzad, Abdus Salam.
■24510▼aEnvironment Generation for Autonomous Agents for Sequential Decision Making
■260 ▼a[Sl]▼bUniversity of California, Berkeley▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a115 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-04, Section: B.
■500 ▼aAdvisor: Stoica, Ion.
■5021 ▼aThesis (Ph.D.)--University of California, Berkeley, 2024.
■520 ▼aAutonomous agents have seen tremendous advancements in solving sequential decision-making problems in recent years, primarily driven by Reinforcement Learning (RL), and more recently, by Large Generative Models. The capability of these autonomous agents depends crucially on the quality and diversity of the learning environments they are trained in. This thesis presents research on designing frameworks and algorithms to formulate and systematically generate environments that improve the generalization capabilities of autonomous agents in solving sequential decision-making tasks. First, we explore the benefits of human-guided programmatic environment generation for training, testing, and debugging autonomous agents in complex real-time strategic (RTS) environments. We present a novel framework that, for the first time, demonstrates the benefits of using scenario specification languages (e.g., SCENIC) for systematic modeling and generation of realistic and diverse RTS RL environments (e.g., Soccer). Next, we discuss a class of algorithms called adaptive teacher Unsupervised Environment Design (UED), which automatically generates training tasks with an RL teacher agent. UED shows promising zero-shot generalization by simultaneously learning a task distribution (i.e., curriculum) and agent policies on the generated tasks. This is a non-stationary process where the task distribution evolves along with agent policies, creating instability over time. While prior works demonstrated the potential of such approaches, training the teacher remained a practical challenge. To this end, we introduce Curriculum Learning via Unsupervised Task Representation Learning (CLUTR): a novel unsupervised curriculum learning algorithm that decouples task representation and curriculum learning into a two-stage optimization to solve the training instability by pretraining a latent task manifold. Following that, we present Multi-Modal Reasoning and Critique for web navigation (MMRC), which introduces augmented environments with multimodal critic agents to enhance the performance of Large Foundational Multimodal Language agents on autonomous web navigation tasks. Together, these approaches portray the importance and usefulness of environment formulation and generation encompassing traditional RL-based and contemporary LLM-based agents.
■590 ▼aSchool code: 0028.
■650 4▼aComputer science
■650 4▼aComputer engineering
■653 ▼aAutonomous agents
■653 ▼aCurriculum learning
■653 ▼aInitiation learning
■653 ▼aLarge language model
■653 ▼aReinforcement learning
■690 ▼a0984
■690 ▼a0464
■690 ▼a0800
■71020▼aUniversity of California, Berkeley▼bComputer Science.
■7730 ▼tDissertations Abstracts International▼g86-04B.
■790 ▼a0028
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17163786▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


