본문

서브메뉴

Advancing Reinforcement Learning: Multi-Agent Optimization, Opportunistic Exploration, and Causal Interpretation
Advancing Reinforcement Learning: Multi-Agent Optimization, Opportunistic Exploration, and...
Advancing Reinforcement Learning: Multi-Agent Optimization, Opportunistic Exploration, and Causal Interpretation

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211150956
ISBN  
9798382604855
DDC  
896
저자명  
Wang, Xiaoxiao.
서명/저자  
Advancing Reinforcement Learning: Multi-Agent Optimization, Opportunistic Exploration, and Causal Interpretation
발행사항  
[Sl] : University of California, Davis, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
204 p
주기사항  
Source: Dissertations Abstracts International, Volume: 85-11, Section: B.
주기사항  
Advisor: Liu, Xin.
학위논문주기  
Thesis (Ph.D.)--University of California, Davis, 2024.
초록/해제  
요약Reinforcement learning (RL), a critical subfield of machine learning, effectively models sequential decision-making scenarios for agents operating in various environments. Despite its extensive applications, RL encounters significant challenges in real world settings, particularly regarding limited data availability, the exploration-exploitation trade-off, and the lack of explainability. In this dissertation, I explore these issues through three distinct lenses. Firstly, I improve data efficiency in situations involving multiple agents or tasks. Secondly, I propose opportunistic learning algorithms in environments with varying exploration costs. Thirdly, I interpret the agent's learned policy through causal explanations.The following sections outline the contributions. Initially, I study online global optimization in multi-agent situations. Cellular network configuration is a suitable application area experiencing these challenges, including the scarcity of diverse historical data, constrained experimental budgets imposed by network operators, and highly complex and unknown network performance functions. To overcome these challenges, I introduce an online-learning-based joint-optimization algorithm combining neural network regression with Gibbs sampling, which considerably outperforms distributed Q-learning in overall performance and ramp-up time. By leveraging similarities among tasks/base stations, I propose a kernel-based multi-task contextual bandit algorithm with the similarity estimated via conditional kernel embedding. These algorithms notably outperform the default cellular network configuration and the respective baseline algorithms.Next, I focus on opportunistic learning, where the exploration cost in RL varies based on different environmental conditions. Given that exploration cost directly impacts the regret of selecting a sub-optimal action, I design the learning strategy to explore more when the cost is low and exploit when the cost is high. I propose an AdaLinUCB algorithm for opportunistic contextual bandits to balance the exploration-exploitation trade-off adaptively. My algorithm significantly outperforms existing contextual bandit algorithms in scenarios with large exploration cost fluctuations. I further develop two algorithms OppUCRL2 and OppPSRL for the finite-horizon episodic Markov decision process, demonstrating the benefits of opportunistic RL. My algorithms balance the exploration-exploitation trade-off dynamically through a variation factor dependent optimism, leading to superior performance. These results are supported by theoretical regret bound analyses ensuring their performance.Lastly, I aim to enhance RL's interpretability by providing causal explanations. My approach quantifies the causal influence of states on actions and their temporal impact, thereby surpassing associative methods in RL policy explanation. I propose a mechanism to quantify the individual-level causal counterfactual path-specific importance score for a structural causal model. This mechanism effectively evaluates causal influence in decision chains, allowing us to comprehend better how a specific decision variable influences an outcome variable.
일반주제명  
African literature
일반주제명  
Computer science
일반주제명  
Electrical engineering
키워드  
Causal explanation
키워드  
Contextual bandit
키워드  
Multi-task learning
키워드  
Opportunistic learning
키워드  
Reinforcement learning
기타저자  
University of California, Davis Computer Science
기본자료저록  
Dissertations Abstracts International. 85-11B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017160317
■00520250211150956
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798382604855
■035    ▼a(MiAaPQ)AAI30993560
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a896
■1001  ▼aWang,  Xiaoxiao.
■24510▼aAdvancing  Reinforcement  Learning:  Multi-Agent  Optimization,  Opportunistic  Exploration,  and  Causal  Interpretation
■260    ▼a[Sl]▼bUniversity  of  California,  Davis▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a204  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  85-11,  Section:  B.
■500    ▼aAdvisor:  Liu,  Xin.
■5021  ▼aThesis  (Ph.D.)--University  of  California,  Davis,  2024.
■520    ▼aReinforcement  learning  (RL),  a  critical  subfield  of  machine  learning,  effectively  models  sequential  decision-making  scenarios  for  agents  operating  in  various  environments.  Despite  its  extensive  applications,  RL  encounters  significant  challenges  in  real  world  settings,  particularly  regarding  limited  data  availability,  the  exploration-exploitation  trade-off,  and  the  lack  of  explainability.  In  this  dissertation,  I  explore  these  issues  through  three  distinct  lenses.  Firstly,  I  improve  data  efficiency  in  situations  involving  multiple  agents  or  tasks.  Secondly,  I  propose  opportunistic  learning  algorithms  in  environments  with  varying  exploration  costs.  Thirdly,  I  interpret  the  agent's  learned  policy  through  causal  explanations.The  following  sections  outline  the  contributions.  Initially,  I  study  online  global  optimization  in  multi-agent  situations.  Cellular  network  configuration  is  a  suitable  application  area  experiencing  these  challenges,  including  the  scarcity  of  diverse  historical  data,  constrained  experimental  budgets  imposed  by  network  operators,  and  highly  complex  and  unknown  network  performance  functions.  To  overcome  these  challenges,  I  introduce  an  online-learning-based  joint-optimization  algorithm  combining  neural  network  regression  with  Gibbs  sampling,  which  considerably  outperforms  distributed  Q-learning  in  overall  performance  and  ramp-up  time.  By  leveraging  similarities  among  tasks/base  stations,  I  propose  a  kernel-based  multi-task  contextual  bandit  algorithm  with  the  similarity  estimated  via  conditional  kernel  embedding.  These  algorithms  notably  outperform  the  default  cellular  network  configuration  and  the  respective  baseline  algorithms.Next,  I  focus  on  opportunistic  learning,  where  the  exploration  cost  in  RL  varies  based  on  different  environmental  conditions.  Given  that  exploration  cost  directly  impacts  the  regret  of  selecting  a  sub-optimal  action,  I  design  the  learning  strategy  to  explore  more  when  the  cost  is  low  and  exploit  when  the  cost  is  high.  I  propose  an  AdaLinUCB  algorithm  for  opportunistic  contextual  bandits  to  balance  the  exploration-exploitation  trade-off  adaptively.  My  algorithm  significantly  outperforms  existing  contextual  bandit  algorithms  in  scenarios  with  large  exploration  cost  fluctuations.  I  further  develop  two  algorithms  OppUCRL2  and  OppPSRL  for  the  finite-horizon  episodic  Markov  decision  process,  demonstrating  the  benefits  of  opportunistic  RL.  My  algorithms  balance  the  exploration-exploitation  trade-off  dynamically  through  a  variation  factor  dependent  optimism,  leading  to  superior  performance.  These  results  are  supported  by  theoretical  regret  bound  analyses  ensuring  their  performance.Lastly,  I  aim  to  enhance  RL's  interpretability  by  providing  causal  explanations.  My  approach  quantifies  the  causal  influence  of  states  on  actions  and  their  temporal  impact,  thereby  surpassing  associative  methods  in  RL  policy  explanation.  I  propose  a  mechanism  to  quantify  the  individual-level  causal  counterfactual  path-specific  importance  score  for  a  structural  causal  model.  This  mechanism  effectively  evaluates  causal  influence  in  decision  chains,  allowing  us  to  comprehend  better  how  a  specific  decision  variable  influences  an  outcome  variable.
■590    ▼aSchool  code:  0029.
■650  4▼aAfrican  literature
■650  4▼aComputer  science
■650  4▼aElectrical  engineering
■653    ▼aCausal  explanation
■653    ▼aContextual  bandit
■653    ▼aMulti-task  learning
■653    ▼aOpportunistic  learning
■653    ▼aReinforcement  learning
■690    ▼a0316
■690    ▼a0984
■690    ▼a0544
■690    ▼a0800
■71020▼aUniversity  of  California,  Davis▼bComputer  Science.
■7730  ▼tDissertations  Abstracts  International▼g85-11B.
■790    ▼a0029
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17160317▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF13914 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.