서브메뉴
검색
Advancing Reinforcement Learning: Multi-Agent Optimization, Opportunistic Exploration, and Causal Interpretation
Advancing Reinforcement Learning: Multi-Agent Optimization, Opportunistic Exploration, and Causal Interpretation
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211150956
- ISBN
- 9798382604855
- DDC
- 896
- 저자명
- Wang, Xiaoxiao.
- 서명/저자
- Advancing Reinforcement Learning: Multi-Agent Optimization, Opportunistic Exploration, and Causal Interpretation
- 발행사항
- [Sl] : University of California, Davis, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 204 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 85-11, Section: B.
- 주기사항
- Advisor: Liu, Xin.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Davis, 2024.
- 초록/해제
- 요약Reinforcement learning (RL), a critical subfield of machine learning, effectively models sequential decision-making scenarios for agents operating in various environments. Despite its extensive applications, RL encounters significant challenges in real world settings, particularly regarding limited data availability, the exploration-exploitation trade-off, and the lack of explainability. In this dissertation, I explore these issues through three distinct lenses. Firstly, I improve data efficiency in situations involving multiple agents or tasks. Secondly, I propose opportunistic learning algorithms in environments with varying exploration costs. Thirdly, I interpret the agent's learned policy through causal explanations.The following sections outline the contributions. Initially, I study online global optimization in multi-agent situations. Cellular network configuration is a suitable application area experiencing these challenges, including the scarcity of diverse historical data, constrained experimental budgets imposed by network operators, and highly complex and unknown network performance functions. To overcome these challenges, I introduce an online-learning-based joint-optimization algorithm combining neural network regression with Gibbs sampling, which considerably outperforms distributed Q-learning in overall performance and ramp-up time. By leveraging similarities among tasks/base stations, I propose a kernel-based multi-task contextual bandit algorithm with the similarity estimated via conditional kernel embedding. These algorithms notably outperform the default cellular network configuration and the respective baseline algorithms.Next, I focus on opportunistic learning, where the exploration cost in RL varies based on different environmental conditions. Given that exploration cost directly impacts the regret of selecting a sub-optimal action, I design the learning strategy to explore more when the cost is low and exploit when the cost is high. I propose an AdaLinUCB algorithm for opportunistic contextual bandits to balance the exploration-exploitation trade-off adaptively. My algorithm significantly outperforms existing contextual bandit algorithms in scenarios with large exploration cost fluctuations. I further develop two algorithms OppUCRL2 and OppPSRL for the finite-horizon episodic Markov decision process, demonstrating the benefits of opportunistic RL. My algorithms balance the exploration-exploitation trade-off dynamically through a variation factor dependent optimism, leading to superior performance. These results are supported by theoretical regret bound analyses ensuring their performance.Lastly, I aim to enhance RL's interpretability by providing causal explanations. My approach quantifies the causal influence of states on actions and their temporal impact, thereby surpassing associative methods in RL policy explanation. I propose a mechanism to quantify the individual-level causal counterfactual path-specific importance score for a structural causal model. This mechanism effectively evaluates causal influence in decision chains, allowing us to comprehend better how a specific decision variable influences an outcome variable.
- 일반주제명
- African literature
- 일반주제명
- Computer science
- 일반주제명
- Electrical engineering
- 기타저자
- University of California, Davis Computer Science
- 기본자료저록
- Dissertations Abstracts International. 85-11B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017160317
■00520250211150956
■006m o d
■007cr#unu||||||||
■020 ▼a9798382604855
■035 ▼a(MiAaPQ)AAI30993560
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a896
■1001 ▼aWang, Xiaoxiao.
■24510▼aAdvancing Reinforcement Learning: Multi-Agent Optimization, Opportunistic Exploration, and Causal Interpretation
■260 ▼a[Sl]▼bUniversity of California, Davis▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a204 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 85-11, Section: B.
■500 ▼aAdvisor: Liu, Xin.
■5021 ▼aThesis (Ph.D.)--University of California, Davis, 2024.
■520 ▼aReinforcement learning (RL), a critical subfield of machine learning, effectively models sequential decision-making scenarios for agents operating in various environments. Despite its extensive applications, RL encounters significant challenges in real world settings, particularly regarding limited data availability, the exploration-exploitation trade-off, and the lack of explainability. In this dissertation, I explore these issues through three distinct lenses. Firstly, I improve data efficiency in situations involving multiple agents or tasks. Secondly, I propose opportunistic learning algorithms in environments with varying exploration costs. Thirdly, I interpret the agent's learned policy through causal explanations.The following sections outline the contributions. Initially, I study online global optimization in multi-agent situations. Cellular network configuration is a suitable application area experiencing these challenges, including the scarcity of diverse historical data, constrained experimental budgets imposed by network operators, and highly complex and unknown network performance functions. To overcome these challenges, I introduce an online-learning-based joint-optimization algorithm combining neural network regression with Gibbs sampling, which considerably outperforms distributed Q-learning in overall performance and ramp-up time. By leveraging similarities among tasks/base stations, I propose a kernel-based multi-task contextual bandit algorithm with the similarity estimated via conditional kernel embedding. These algorithms notably outperform the default cellular network configuration and the respective baseline algorithms.Next, I focus on opportunistic learning, where the exploration cost in RL varies based on different environmental conditions. Given that exploration cost directly impacts the regret of selecting a sub-optimal action, I design the learning strategy to explore more when the cost is low and exploit when the cost is high. I propose an AdaLinUCB algorithm for opportunistic contextual bandits to balance the exploration-exploitation trade-off adaptively. My algorithm significantly outperforms existing contextual bandit algorithms in scenarios with large exploration cost fluctuations. I further develop two algorithms OppUCRL2 and OppPSRL for the finite-horizon episodic Markov decision process, demonstrating the benefits of opportunistic RL. My algorithms balance the exploration-exploitation trade-off dynamically through a variation factor dependent optimism, leading to superior performance. These results are supported by theoretical regret bound analyses ensuring their performance.Lastly, I aim to enhance RL's interpretability by providing causal explanations. My approach quantifies the causal influence of states on actions and their temporal impact, thereby surpassing associative methods in RL policy explanation. I propose a mechanism to quantify the individual-level causal counterfactual path-specific importance score for a structural causal model. This mechanism effectively evaluates causal influence in decision chains, allowing us to comprehend better how a specific decision variable influences an outcome variable.
■590 ▼aSchool code: 0029.
■650 4▼aAfrican literature
■650 4▼aComputer science
■650 4▼aElectrical engineering
■653 ▼aCausal explanation
■653 ▼aContextual bandit
■653 ▼aMulti-task learning
■653 ▼aOpportunistic learning
■653 ▼aReinforcement learning
■690 ▼a0316
■690 ▼a0984
■690 ▼a0544
■690 ▼a0800
■71020▼aUniversity of California, Davis▼bComputer Science.
■7730 ▼tDissertations Abstracts International▼g85-11B.
■790 ▼a0029
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17160317▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


