본문

서브메뉴

Environment Generation for Autonomous Agents for Sequential Decision Making
Environment Generation for Autonomous Agents for Sequential Decision Making
Environment Generation for Autonomous Agents for Sequential Decision Making

Detailed Information

자료유형  
 학위논문 서양
최종처리일시  
20250211152753
ISBN  
9798384448891
DDC  
004
저자명  
Azad, Abdus Salam.
서명/저자  
Environment Generation for Autonomous Agents for Sequential Decision Making
발행사항  
[Sl] : University of California, Berkeley, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
115 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-04, Section: B.
주기사항  
Advisor: Stoica, Ion.
학위논문주기  
Thesis (Ph.D.)--University of California, Berkeley, 2024.
초록/해제  
요약Autonomous agents have seen tremendous advancements in solving sequential decision-making problems in recent years, primarily driven by Reinforcement Learning (RL), and more recently, by Large Generative Models. The capability of these autonomous agents depends crucially on the quality and diversity of the learning environments they are trained in. This thesis presents research on designing frameworks and algorithms to formulate and systematically generate environments that improve the generalization capabilities of autonomous agents in solving sequential decision-making tasks. First, we explore the benefits of human-guided programmatic environment generation for training, testing, and debugging autonomous agents in complex real-time strategic (RTS) environments. We present a novel framework that, for the first time, demonstrates the benefits of using scenario specification languages (e.g., SCENIC) for systematic modeling and generation of realistic and diverse RTS RL environments (e.g., Soccer). Next, we discuss a class of algorithms called adaptive teacher Unsupervised Environment Design (UED), which automatically generates training tasks with an RL teacher agent. UED shows promising zero-shot generalization by simultaneously learning a task distribution (i.e., curriculum) and agent policies on the generated tasks. This is a non-stationary process where the task distribution evolves along with agent policies, creating instability over time. While prior works demonstrated the potential of such approaches, training the teacher remained a practical challenge. To this end, we introduce Curriculum Learning via Unsupervised Task Representation Learning (CLUTR): a novel unsupervised curriculum learning algorithm that decouples task representation and curriculum learning into a two-stage optimization to solve the training instability by pretraining a latent task manifold. Following that, we present Multi-Modal Reasoning and Critique for web navigation (MMRC), which introduces augmented environments with multimodal critic agents to enhance the performance of Large Foundational Multimodal Language agents on autonomous web navigation tasks. Together, these approaches portray the importance and usefulness of environment formulation and generation encompassing traditional RL-based and contemporary LLM-based agents.
일반주제명  
Computer science
일반주제명  
Computer engineering
키워드  
Autonomous agents
키워드  
Curriculum learning
키워드  
Initiation learning
키워드  
Large language model
키워드  
Reinforcement learning
기타저자  
University of California, Berkeley Computer Science
기본자료저록  
Dissertations Abstracts International. 86-04B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017163786
■00520250211152753
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798384448891
■035    ▼a(MiAaPQ)AAI31555572
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aAzad,  Abdus  Salam.
■24510▼aEnvironment  Generation  for  Autonomous  Agents  for  Sequential  Decision  Making
■260    ▼a[Sl]▼bUniversity  of  California,  Berkeley▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a115  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-04,  Section:  B.
■500    ▼aAdvisor:  Stoica,  Ion.
■5021  ▼aThesis  (Ph.D.)--University  of  California,  Berkeley,  2024.
■520    ▼aAutonomous  agents  have  seen  tremendous  advancements  in  solving  sequential  decision-making  problems  in  recent  years,  primarily  driven  by  Reinforcement  Learning  (RL),  and  more  recently,  by  Large  Generative  Models.  The  capability  of  these  autonomous  agents  depends  crucially  on  the  quality  and  diversity  of  the  learning  environments  they  are  trained  in.  This  thesis  presents  research  on  designing  frameworks  and  algorithms  to  formulate  and  systematically  generate  environments  that  improve  the  generalization  capabilities  of  autonomous  agents  in  solving  sequential  decision-making  tasks.  First,  we  explore  the  benefits  of  human-guided  programmatic  environment  generation  for  training,  testing,  and  debugging  autonomous  agents  in  complex  real-time  strategic  (RTS)  environments.  We  present  a  novel  framework  that,  for  the  first  time,  demonstrates  the  benefits  of  using  scenario  specification  languages  (e.g.,  SCENIC)  for  systematic  modeling  and  generation  of  realistic  and  diverse  RTS  RL  environments  (e.g.,  Soccer).  Next,  we  discuss  a  class  of  algorithms  called  adaptive  teacher  Unsupervised  Environment  Design  (UED),  which  automatically  generates  training  tasks  with  an  RL  teacher  agent.  UED  shows  promising  zero-shot  generalization  by  simultaneously  learning  a  task  distribution  (i.e.,  curriculum)  and  agent  policies  on  the  generated  tasks.  This  is  a  non-stationary  process  where  the  task  distribution  evolves  along  with  agent  policies,  creating  instability  over  time.  While  prior  works  demonstrated  the  potential  of  such  approaches,  training  the  teacher  remained  a  practical  challenge.  To  this  end,  we  introduce  Curriculum  Learning  via  Unsupervised  Task  Representation  Learning  (CLUTR):  a  novel  unsupervised  curriculum  learning  algorithm  that  decouples  task  representation  and  curriculum  learning  into  a  two-stage  optimization  to  solve  the  training  instability  by  pretraining  a  latent  task  manifold.  Following  that,  we  present  Multi-Modal  Reasoning  and  Critique  for  web  navigation  (MMRC),  which  introduces  augmented  environments  with  multimodal  critic  agents  to  enhance  the  performance  of  Large  Foundational  Multimodal  Language  agents  on  autonomous  web  navigation  tasks.  Together,  these  approaches  portray  the  importance  and  usefulness  of  environment  formulation  and  generation  encompassing  traditional  RL-based  and  contemporary  LLM-based  agents.
■590    ▼aSchool  code:  0028.
■650  4▼aComputer  science
■650  4▼aComputer  engineering
■653    ▼aAutonomous  agents
■653    ▼aCurriculum  learning
■653    ▼aInitiation  learning
■653    ▼aLarge  language  model
■653    ▼aReinforcement  learning
■690    ▼a0984
■690    ▼a0464
■690    ▼a0800
■71020▼aUniversity  of  California,  Berkeley▼bComputer  Science.
■7730  ▼tDissertations  Abstracts  International▼g86-04B.
■790    ▼a0028
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17163786▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

Preview

Export

ChatGPT Discussion

AI Recommended Related Books


    New Books MORE
    Statistics for the past 3 years. Go to brief

    Подробнее информация.

    • Бронирование
    • не существует
    • моя папка
    • Первый запрос зрения
    • Non-Book Loan Application
    • Nighttime Book Loan Application
    материал
    Reg No. Количество платежных Местоположение статус Ленд информации
    TF13679 전자도서 대출가능 My Folder 부재도서신고 비도서대출신청 야간 도서대출신청

    * Бронирование доступны в заимствований книги. Чтобы сделать предварительный заказ, пожалуйста, нажмите кнопку бронирование

    Books borrowed together with this book

    Related Popular Books

    Available after logging in.