본문

서브메뉴

Reinforcement Learning-Based Human Operator Decision Support Agent for Highly Transient Industrial Processes
Reinforcement Learning-Based Human Operator Decision Support Agent for Highly Transient In...
Reinforcement Learning-Based Human Operator Decision Support Agent for Highly Transient Industrial Processes

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211151408
ISBN  
9798342102131
DDC  
515
저자명  
Ruan, Jianqi.
서명/저자  
Reinforcement Learning-Based Human Operator Decision Support Agent for Highly Transient Industrial Processes
발행사항  
[Sl] : Purdue University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
103 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-04, Section: B.
주기사항  
Advisor: Jain, Neera.
학위논문주기  
Thesis (Ph.D.)--Purdue University, 2024.
초록/해제  
요약Most industrial processes are not fully-automated. Although reference tracking can be handled by low-level controllers, initializing and adjusting the reference, or setpoint, values, are commonly tasks assigned to human operators. A major challenge that arises, though, is control policy variation among operators which in turn results in inconsistencies in the final product. In order to guide operators to pursue better and more consistent performance, researchers have explored the optimal control policy through different approaches. Although in different applications, researchers use different approaches, an accurate process model is still crucial to the approaches. However, for a highly transient process (e.g., the startup of a manufacturing process), modeling can be challenging and inaccurate, and approaches highly relying on a process model may not work well. One example is process startup in a twin-roll steel strip casting process and motivates this work.In this dissertation, I propose three offline reinforcement learning (RL) algorithms which require the RL agent to learn a control policy from a fixed dataset that is pre-collected by human operators during operations of the twin-roll casting process. Compared to existing offline RL algorithms, the proposed algorithms focus on exploiting the best control policy used by human operators rather than exploring new control policies constrained by the existing policies. In addition, in existing offline RL algorithms, there is not enough consideration of the imbalanced dataset problem. In the second and the third proposed algorithms, I leverage the idea of cost sensitive learning to incentivize the RL agent to learn the most valuable control policy, rather than the most common one represented in the dataset. In addition, since the process model is not available, I propose a performance metric that does not require a process model or simulator for agent testing. The third proposed algorithm is compared with benchmark offline RL algorithms and achieves better and more consistent performance.
일반주제명  
Dynamical systems
일반주제명  
Decision making
일반주제명  
Neural networks
일반주제명  
Mathematics
일반주제명  
Statistics
기타저자  
Purdue University.
기본자료저록  
Dissertations Abstracts International. 86-04B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017161524
■00520250211151408
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798342102131
■035    ▼a(MiAaPQ)AAI31285273
■035    ▼a(MiAaPQ)25302148
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a515
■1001  ▼aRuan,  Jianqi.
■24510▼aReinforcement  Learning-Based  Human  Operator  Decision  Support  Agent  for  Highly  Transient  Industrial  Processes
■260    ▼a[Sl]▼bPurdue  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a103  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-04,  Section:  B.
■500    ▼aAdvisor:  Jain,  Neera.
■5021  ▼aThesis  (Ph.D.)--Purdue  University,  2024.
■520    ▼aMost  industrial  processes  are  not  fully-automated.  Although  reference  tracking  can  be  handled  by  low-level  controllers,  initializing  and  adjusting  the  reference,  or  setpoint,  values,  are  commonly  tasks  assigned  to  human  operators.  A  major  challenge  that  arises,  though,  is  control  policy  variation  among  operators  which  in  turn  results  in  inconsistencies  in  the  final  product.  In  order  to  guide  operators  to  pursue  better  and  more  consistent  performance,  researchers  have  explored  the  optimal  control  policy  through  different  approaches.  Although  in  different  applications,  researchers  use  different  approaches,  an  accurate  process  model  is  still  crucial  to  the  approaches.  However,  for  a  highly  transient  process  (e.g.,  the  startup  of  a  manufacturing  process),  modeling  can  be  challenging  and  inaccurate,  and  approaches  highly  relying  on  a  process  model  may  not  work  well.  One  example  is  process  startup  in  a  twin-roll  steel  strip  casting  process  and  motivates  this  work.In  this  dissertation,  I  propose  three  offline  reinforcement  learning  (RL)  algorithms  which  require  the  RL  agent  to  learn  a  control  policy  from  a  fixed  dataset  that  is  pre-collected  by  human  operators  during  operations  of  the  twin-roll  casting  process.  Compared  to  existing  offline  RL  algorithms,  the  proposed  algorithms  focus  on  exploiting  the  best  control  policy  used  by  human  operators  rather  than  exploring  new  control  policies  constrained  by  the  existing  policies.  In  addition,  in  existing  offline  RL  algorithms,  there  is  not  enough  consideration  of  the  imbalanced  dataset  problem.  In  the  second  and  the  third  proposed  algorithms,  I  leverage  the  idea  of  cost  sensitive  learning  to  incentivize  the  RL  agent  to  learn  the  most  valuable  control  policy,  rather  than  the  most  common  one  represented  in  the  dataset.  In  addition,  since  the  process  model  is  not  available,  I  propose  a  performance  metric  that  does  not  require  a  process  model  or  simulator  for  agent  testing.  The  third  proposed  algorithm  is  compared  with  benchmark  offline  RL  algorithms  and  achieves  better  and  more  consistent  performance.
■590    ▼aSchool  code:  0183.
■650  4▼aDynamical  systems
■650  4▼aDecision  making
■650  4▼aNeural  networks
■650  4▼aMathematics
■650  4▼aStatistics
■690    ▼a0800
■690    ▼a0405
■690    ▼a0463
■71020▼aPurdue  University.
■7730  ▼tDissertations  Abstracts  International▼g86-04B.
■790    ▼a0183
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17161524▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF09924 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.