서브메뉴
검색
Towards Specialized Reinforcement Learning From Diverse Data
Towards Specialized Reinforcement Learning From Diverse Data
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211152008
- ISBN
- 9798384048411
- DDC
- 004
- 서명/저자
- Towards Specialized Reinforcement Learning From Diverse Data
- 발행사항
- [Sl] : Cornell University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 337 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-03, Section: B.
- 주기사항
- Advisor: Sun, Wen.
- 학위논문주기
- Thesis (Ph.D.)--Cornell University, 2024.
- 초록/해제
- 요약Reinforcement learning (RL) fundamentally focuses on teaching agents how to make decisions by interacting with an environment. Unlike supervised learning approaches that learn from a fixed dataset, reinforcement learning agents learn by doing, receiving feedback as rewards or penalties based on their chosen actions. However, the generality of RL makes efficient learning incredibly difficult, making adopting novel tasks complicated. Even with the explosive success of ChatGPT (OpenAI, 2023), which applied a deep RL algorithm, Proximal Policy Optimization (PPO) (Schulman et al., 2017b), to large language models (LLMs), there has been more interest in either reducing the complexity of RL algorithms or eliminating the need for online interaction. Even beyond LLMs, despite RL's superhuman abilities in games such as DOTA (Berner et al., 2019) or Go (Silver et al., 2016b), we have yet to see widespread adoption of RL in real-world applications compared to other, arguably more specialized, learning paradigms such as supervised learning. We notice that a critical challenge for the broad adoption of RL is efficiently utilizing diverse data sources to create a specialized algorithm. That is, for many of the successes mentioned above in RL, the algorithms used to learn the agents were general-purpose algorithms that could also be used in other applications.In this thesis, we attempt to introduce RL algorithms that progress toward specialized algorithms for various settings. We first discuss doing efficient inverse reinforcement learning (IRL) from different types of data sources. We consider three settings: learning from observations alone, where the demonstration data does not contain action information; offline learning, where instead of interactive access to the environment we only get a large dataset of interactions; and off-policy learning where the interactive feedback that we learn from can be from different learning agents. In all three settings, we introduce a principled algorithm that performs efficient learning in a wide range of control tasks. Next we discuss learning specialized algorithms in the space of generative models. Foundation models increasingly live up to their namesake, becoming capable base models for improved downstream performance on various tasks across multiple application domains. In this thesis's second part, we investigate RL with these models for learning decision-making agents from diverse data sources. We present three different learning settings with generative models: text generation with an interactive black box model, text generation with high-quality human labels, and text instruction-guided image generation. Overall, each setting has a specific property, whether it is deterministic transition dynamics or a short horizon, that allows for the design of more specialized algorithms that efficiently exploit these properties and improves beyond the general RL baseline.
- 일반주제명
- Computer science
- 키워드
- Offline learning
- 기타저자
- Cornell University Computer Science
- 기본자료저록
- Dissertations Abstracts International. 86-03B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017162399
■00520250211152008
■006m o d
■007cr#unu||||||||
■020 ▼a9798384048411
■035 ▼a(MiAaPQ)AAI31330820
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aChang, Jonathan Daniel.▼0(orcid)0009-0007-5038-0805
■24510▼aTowards Specialized Reinforcement Learning From Diverse Data
■260 ▼a[Sl]▼bCornell University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a337 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-03, Section: B.
■500 ▼aAdvisor: Sun, Wen.
■5021 ▼aThesis (Ph.D.)--Cornell University, 2024.
■520 ▼aReinforcement learning (RL) fundamentally focuses on teaching agents how to make decisions by interacting with an environment. Unlike supervised learning approaches that learn from a fixed dataset, reinforcement learning agents learn by doing, receiving feedback as rewards or penalties based on their chosen actions. However, the generality of RL makes efficient learning incredibly difficult, making adopting novel tasks complicated. Even with the explosive success of ChatGPT (OpenAI, 2023), which applied a deep RL algorithm, Proximal Policy Optimization (PPO) (Schulman et al., 2017b), to large language models (LLMs), there has been more interest in either reducing the complexity of RL algorithms or eliminating the need for online interaction. Even beyond LLMs, despite RL's superhuman abilities in games such as DOTA (Berner et al., 2019) or Go (Silver et al., 2016b), we have yet to see widespread adoption of RL in real-world applications compared to other, arguably more specialized, learning paradigms such as supervised learning. We notice that a critical challenge for the broad adoption of RL is efficiently utilizing diverse data sources to create a specialized algorithm. That is, for many of the successes mentioned above in RL, the algorithms used to learn the agents were general-purpose algorithms that could also be used in other applications.In this thesis, we attempt to introduce RL algorithms that progress toward specialized algorithms for various settings. We first discuss doing efficient inverse reinforcement learning (IRL) from different types of data sources. We consider three settings: learning from observations alone, where the demonstration data does not contain action information; offline learning, where instead of interactive access to the environment we only get a large dataset of interactions; and off-policy learning where the interactive feedback that we learn from can be from different learning agents. In all three settings, we introduce a principled algorithm that performs efficient learning in a wide range of control tasks. Next we discuss learning specialized algorithms in the space of generative models. Foundation models increasingly live up to their namesake, becoming capable base models for improved downstream performance on various tasks across multiple application domains. In this thesis's second part, we investigate RL with these models for learning decision-making agents from diverse data sources. We present three different learning settings with generative models: text generation with an interactive black box model, text generation with high-quality human labels, and text instruction-guided image generation. Overall, each setting has a specific property, whether it is deterministic transition dynamics or a short horizon, that allows for the design of more specialized algorithms that efficiently exploit these properties and improves beyond the general RL baseline.
■590 ▼aSchool code: 0058.
■650 4▼aComputer science
■653 ▼aGenerative models
■653 ▼aImitation learning
■653 ▼aOffline learning
■653 ▼aReinforcement learning
■653 ▼aProximal Policy Optimization
■690 ▼a0984
■690 ▼a0800
■71020▼aCornell University▼bComputer Science.
■7730 ▼tDissertations Abstracts International▼g86-03B.
■790 ▼a0058
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17162399▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


