본문

서브메뉴

Towards Specialized Reinforcement Learning From Diverse Data
Towards Specialized Reinforcement Learning From Diverse Data
Towards Specialized Reinforcement Learning From Diverse Data

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211152008
ISBN  
9798384048411
DDC  
004
저자명  
Chang, Jonathan Daniel.
서명/저자  
Towards Specialized Reinforcement Learning From Diverse Data
발행사항  
[Sl] : Cornell University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
337 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-03, Section: B.
주기사항  
Advisor: Sun, Wen.
학위논문주기  
Thesis (Ph.D.)--Cornell University, 2024.
초록/해제  
요약Reinforcement learning (RL) fundamentally focuses on teaching agents how to make decisions by interacting with an environment. Unlike supervised learning approaches that learn from a fixed dataset, reinforcement learning agents learn by doing, receiving feedback as rewards or penalties based on their chosen actions. However, the generality of RL makes efficient learning incredibly difficult, making adopting novel tasks complicated. Even with the explosive success of ChatGPT (OpenAI, 2023), which applied a deep RL algorithm, Proximal Policy Optimization (PPO) (Schulman et al., 2017b), to large language models (LLMs), there has been more interest in either reducing the complexity of RL algorithms or eliminating the need for online interaction. Even beyond LLMs, despite RL's superhuman abilities in games such as DOTA (Berner et al., 2019) or Go (Silver et al., 2016b), we have yet to see widespread adoption of RL in real-world applications compared to other, arguably more specialized, learning paradigms such as supervised learning. We notice that a critical challenge for the broad adoption of RL is efficiently utilizing diverse data sources to create a specialized algorithm. That is, for many of the successes mentioned above in RL, the algorithms used to learn the agents were general-purpose algorithms that could also be used in other applications.In this thesis, we attempt to introduce RL algorithms that progress toward specialized algorithms for various settings. We first discuss doing efficient inverse reinforcement learning (IRL) from different types of data sources. We consider three settings: learning from observations alone, where the demonstration data does not contain action information; offline learning, where instead of interactive access to the environment we only get a large dataset of interactions; and off-policy learning where the interactive feedback that we learn from can be from different learning agents. In all three settings, we introduce a principled algorithm that performs efficient learning in a wide range of control tasks. Next we discuss learning specialized algorithms in the space of generative models. Foundation models increasingly live up to their namesake, becoming capable base models for improved downstream performance on various tasks across multiple application domains. In this thesis's second part, we investigate RL with these models for learning decision-making agents from diverse data sources. We present three different learning settings with generative models: text generation with an interactive black box model, text generation with high-quality human labels, and text instruction-guided image generation. Overall, each setting has a specific property, whether it is deterministic transition dynamics or a short horizon, that allows for the design of more specialized algorithms that efficiently exploit these properties and improves beyond the general RL baseline.
일반주제명  
Computer science
키워드  
Generative models
키워드  
Imitation learning
키워드  
Offline learning
키워드  
Reinforcement learning
키워드  
Proximal Policy Optimization
기타저자  
Cornell University Computer Science
기본자료저록  
Dissertations Abstracts International. 86-03B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017162399
■00520250211152008
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798384048411
■035    ▼a(MiAaPQ)AAI31330820
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aChang,  Jonathan  Daniel.▼0(orcid)0009-0007-5038-0805
■24510▼aTowards  Specialized  Reinforcement  Learning  From  Diverse  Data
■260    ▼a[Sl]▼bCornell  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a337  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-03,  Section:  B.
■500    ▼aAdvisor:  Sun,  Wen.
■5021  ▼aThesis  (Ph.D.)--Cornell  University,  2024.
■520    ▼aReinforcement  learning  (RL)  fundamentally  focuses  on  teaching  agents  how  to  make  decisions  by  interacting  with  an  environment.  Unlike  supervised  learning  approaches  that  learn  from  a  fixed  dataset,  reinforcement  learning  agents  learn  by  doing,  receiving  feedback  as  rewards  or  penalties  based  on  their  chosen  actions.  However,  the  generality  of  RL  makes  efficient  learning  incredibly  difficult,  making  adopting  novel  tasks  complicated.  Even  with  the  explosive  success  of  ChatGPT  (OpenAI,  2023),  which  applied  a  deep  RL  algorithm,  Proximal  Policy  Optimization  (PPO)  (Schulman  et  al.,  2017b),  to  large  language  models  (LLMs),  there  has  been  more  interest  in  either  reducing  the  complexity  of  RL  algorithms  or  eliminating  the  need  for  online  interaction.  Even  beyond  LLMs,  despite  RL's  superhuman  abilities  in  games  such  as  DOTA  (Berner  et  al.,  2019)  or  Go  (Silver  et  al.,  2016b),  we  have  yet  to  see  widespread  adoption  of  RL  in  real-world  applications  compared  to  other,  arguably  more  specialized,  learning  paradigms  such  as  supervised  learning.  We  notice  that  a  critical  challenge  for  the  broad  adoption  of  RL  is  efficiently  utilizing  diverse  data  sources  to  create  a  specialized  algorithm.  That  is,  for  many  of  the  successes  mentioned  above  in  RL,  the  algorithms  used  to  learn  the  agents  were  general-purpose  algorithms  that  could  also  be  used  in  other  applications.In  this  thesis,  we  attempt  to  introduce  RL  algorithms  that  progress  toward  specialized  algorithms  for  various  settings.  We  first  discuss  doing  efficient  inverse  reinforcement  learning  (IRL)  from  different  types  of  data  sources.  We  consider  three  settings:  learning  from  observations  alone,  where  the  demonstration  data  does  not  contain  action  information;  offline  learning,  where  instead  of  interactive  access  to  the  environment  we  only  get  a  large  dataset  of  interactions;  and  off-policy  learning  where  the  interactive  feedback  that  we  learn  from  can  be  from  different  learning  agents.  In  all  three  settings,  we  introduce  a  principled  algorithm  that  performs  efficient  learning  in  a  wide  range  of  control  tasks.  Next  we  discuss  learning  specialized  algorithms  in  the  space  of  generative  models.  Foundation  models  increasingly  live  up  to  their  namesake,  becoming  capable  base  models  for  improved  downstream  performance  on  various  tasks  across  multiple  application  domains.  In  this  thesis's  second  part,  we  investigate  RL  with  these  models  for  learning  decision-making  agents  from  diverse  data  sources.  We  present  three  different  learning  settings  with  generative  models:  text  generation  with  an  interactive  black  box  model,  text  generation  with  high-quality  human  labels,  and  text  instruction-guided  image  generation.  Overall,  each  setting  has  a  specific  property,  whether  it  is  deterministic  transition  dynamics  or  a  short  horizon,  that  allows  for  the  design  of  more  specialized  algorithms  that  efficiently  exploit  these  properties  and  improves  beyond  the  general  RL  baseline.
■590    ▼aSchool  code:  0058.
■650  4▼aComputer  science
■653    ▼aGenerative  models
■653    ▼aImitation  learning
■653    ▼aOffline  learning
■653    ▼aReinforcement  learning
■653    ▼aProximal  Policy  Optimization
■690    ▼a0984
■690    ▼a0800
■71020▼aCornell  University▼bComputer  Science.
■7730  ▼tDissertations  Abstracts  International▼g86-03B.
■790    ▼a0058
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17162399▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF13835 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.