본문

서브메뉴

Learning From Human Feedback: Ranking, Bandit, and Preference Optimization
Learning From Human Feedback: Ranking, Bandit, and Preference Optimization
Learning From Human Feedback: Ranking, Bandit, and Preference Optimization

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211153146
ISBN  
9798384464778
DDC  
004
저자명  
Wu, Yue.
서명/저자  
Learning From Human Feedback: Ranking, Bandit, and Preference Optimization
발행사항  
[Sl] : University of California, Los Angeles, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
188 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-04, Section: A.
주기사항  
Advisor: Gu, Quanquan.
학위논문주기  
Thesis (Ph.D.)--University of California, Los Angeles, 2024.
초록/해제  
요약This dissertation investigates several challenges in artificial intelligence (AI) alignment and reinforcement learning (RL), particularly focusing on applications when only preference feedback is available. Learning from preference feedback has been one central problem across different fields such as ranking, recommendation systems, and social choice theory. Recently, reinforcement learning from human feedback (RLHF) has also shown its strong potential in utilizing weakly supervised human data (preference feedback) and its ability to encode human values into machine learning models accurately. This dissertation aims to comprehensively characterize preference-based statistical learning, focusing on the sample complexity of ranking and preference model estimation and fine-tuning large language models.The first part of the dissertation explores novel methods for the learning-to-rank problem. I studied learning to rank under the strong stochastic transitivity (SST) condition, a prevalent model without assuming a score for each option. SST assumes that the accuracy of the comparison between two items increases as the disparity in their quality widens. I proposed one of the first adaptive approaches that can effectively aggregate the feedback from different human labelers and illustrated how the relationship between the number of human queries and resulting performance depends on the properties of the human labelers. I further follow up in this direction with my collaborators and provide active ranking algorithms that can work without strong stochastic transitivity. We developed novel algorithms under this practical yet harder setting. Our efficient algorithm requires fewer human queries compared with algorithms designed with stronger assumptions. The algorithm is provably optimal. The second part focuses on preference learning without transitivity assumption. In reality, humans rarely make consistent comparisons and often demonstrate contradicting preferences such as a loop within the preference relations. I considered the most general setting where there is no transitivity at all. I proposed algorithms that identify the Borda winner, an optimal choice even when a true underlying rank does not exist. I showed the algorithm enjoys minimum regret, a notion that trades between exploration and exploitation. This result sheds light on the fundamental difficulty and cost of recovering human preferences under the fewest assumptions.The third part also focuses on preference learning without transitivity assumption, but instead considers an alternative definition of learning objective, the von Neumann winner. I first formulate the general preference as a game environment where two players aim to win over each other and then present an algorithmic framework that can solve this game in a self-play manner asymptotically. The algorithm is then extended to the task of fine-tuning large language models and shows remarkable empirical performance.The methods and techniques discussed in this dissertation cover a full spectrum of different assumptions and settings of preference learning. In each setting, the new algorithms are presented along with theoretical analysis ensuring a tight performance guarantee. Additionally, during the exploration of different settings, new research directions and open questions are identified, which could help promote the research of preference learning in terms of sample efficiency in the future.
일반주제명  
Computer science
일반주제명  
Statistics
일반주제명  
Information science
키워드  
Bandits
키워드  
Human feedback
키워드  
Ranking
키워드  
Reinforcement learning
키워드  
Preference optimization
기타저자  
University of California, Los Angeles Computer Science 0201
기본자료저록  
Dissertations Abstracts International. 86-04A.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017165186
■00520250211153146
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798384464778
■035    ▼a(MiAaPQ)AAI31563320
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aWu,  Yue.
■24510▼aLearning  From  Human  Feedback:  Ranking,  Bandit,  and  Preference  Optimization
■260    ▼a[Sl]▼bUniversity  of  California,  Los  Angeles▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a188  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-04,  Section:  A.
■500    ▼aAdvisor:  Gu,  Quanquan.
■5021  ▼aThesis  (Ph.D.)--University  of  California,  Los  Angeles,  2024.
■520    ▼aThis  dissertation  investigates  several  challenges  in  artificial  intelligence  (AI)  alignment  and  reinforcement  learning  (RL),  particularly  focusing  on  applications  when  only  preference  feedback  is  available.  Learning  from  preference  feedback  has  been  one  central  problem  across  different  fields  such  as  ranking,  recommendation  systems,  and  social  choice  theory.  Recently,  reinforcement  learning  from  human  feedback  (RLHF)  has  also  shown  its  strong  potential  in  utilizing  weakly  supervised  human  data  (preference  feedback)  and  its  ability  to  encode  human  values  into  machine  learning  models  accurately.  This  dissertation  aims  to  comprehensively  characterize  preference-based  statistical  learning,  focusing  on  the  sample  complexity  of  ranking  and  preference  model  estimation  and  fine-tuning  large  language  models.The  first  part  of  the  dissertation  explores  novel  methods  for  the  learning-to-rank  problem.  I  studied  learning  to  rank  under  the  strong  stochastic  transitivity  (SST)  condition,  a  prevalent  model  without  assuming  a  score  for  each  option.  SST  assumes  that  the  accuracy  of  the  comparison  between  two  items  increases  as  the  disparity  in  their  quality  widens.  I  proposed  one  of  the  first  adaptive  approaches  that  can  effectively  aggregate  the  feedback  from  different  human  labelers  and  illustrated  how  the  relationship  between  the  number  of  human  queries  and  resulting  performance  depends  on  the  properties  of  the  human  labelers.  I  further  follow  up  in  this  direction  with  my  collaborators  and  provide  active  ranking  algorithms  that  can  work  without  strong  stochastic  transitivity.  We  developed  novel  algorithms  under  this  practical  yet  harder  setting.  Our  efficient  algorithm  requires  fewer  human  queries  compared  with  algorithms  designed  with  stronger  assumptions.  The  algorithm  is  provably  optimal.  The  second  part  focuses  on  preference  learning  without  transitivity  assumption.  In  reality,  humans  rarely  make  consistent  comparisons  and  often  demonstrate  contradicting  preferences  such  as  a  loop  within  the  preference  relations.  I  considered  the  most  general  setting  where  there  is  no  transitivity  at  all.  I  proposed  algorithms  that  identify  the  Borda  winner,  an  optimal  choice  even  when  a  true  underlying  rank  does  not  exist.  I  showed  the  algorithm  enjoys  minimum  regret,  a  notion  that  trades  between  exploration  and  exploitation.  This  result  sheds  light  on  the  fundamental  difficulty  and  cost  of  recovering  human  preferences  under  the  fewest  assumptions.The  third  part  also  focuses  on  preference  learning  without  transitivity  assumption,  but  instead  considers  an  alternative  definition  of  learning  objective,  the  von  Neumann  winner.  I  first  formulate  the  general  preference  as  a  game  environment  where  two  players  aim  to  win  over  each  other  and  then  present  an  algorithmic  framework  that  can  solve  this  game  in  a  self-play  manner  asymptotically.  The  algorithm  is  then  extended  to  the  task  of  fine-tuning  large  language  models  and  shows  remarkable  empirical  performance.The  methods  and  techniques  discussed  in  this  dissertation  cover  a  full  spectrum  of  different  assumptions  and  settings  of  preference  learning.  In  each  setting,  the  new  algorithms  are  presented  along  with  theoretical  analysis  ensuring  a  tight  performance  guarantee.  Additionally,  during  the  exploration  of  different  settings,  new  research  directions  and  open  questions  are  identified,  which  could  help  promote  the  research  of  preference  learning  in  terms  of  sample  efficiency  in  the  future.
■590    ▼aSchool  code:  0031.
■650  4▼aComputer  science
■650  4▼aStatistics
■650  4▼aInformation  science
■653    ▼aBandits
■653    ▼aHuman  feedback
■653    ▼aRanking
■653    ▼aReinforcement  learning
■653    ▼aPreference  optimization
■690    ▼a0800
■690    ▼a0984
■690    ▼a0723
■690    ▼a0463
■71020▼aUniversity  of  California,  Los  Angeles▼bComputer  Science  0201.
■7730  ▼tDissertations  Abstracts  International▼g86-04A.
■790    ▼a0031
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17165186▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF13485 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.