서브메뉴
검색
Learning From Human Feedback: Ranking, Bandit, and Preference Optimization
Learning From Human Feedback: Ranking, Bandit, and Preference Optimization
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211153146
- ISBN
- 9798384464778
- DDC
- 004
- 저자명
- Wu, Yue.
- 서명/저자
- Learning From Human Feedback: Ranking, Bandit, and Preference Optimization
- 발행사항
- [Sl] : University of California, Los Angeles, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 188 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-04, Section: A.
- 주기사항
- Advisor: Gu, Quanquan.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Los Angeles, 2024.
- 초록/해제
- 요약This dissertation investigates several challenges in artificial intelligence (AI) alignment and reinforcement learning (RL), particularly focusing on applications when only preference feedback is available. Learning from preference feedback has been one central problem across different fields such as ranking, recommendation systems, and social choice theory. Recently, reinforcement learning from human feedback (RLHF) has also shown its strong potential in utilizing weakly supervised human data (preference feedback) and its ability to encode human values into machine learning models accurately. This dissertation aims to comprehensively characterize preference-based statistical learning, focusing on the sample complexity of ranking and preference model estimation and fine-tuning large language models.The first part of the dissertation explores novel methods for the learning-to-rank problem. I studied learning to rank under the strong stochastic transitivity (SST) condition, a prevalent model without assuming a score for each option. SST assumes that the accuracy of the comparison between two items increases as the disparity in their quality widens. I proposed one of the first adaptive approaches that can effectively aggregate the feedback from different human labelers and illustrated how the relationship between the number of human queries and resulting performance depends on the properties of the human labelers. I further follow up in this direction with my collaborators and provide active ranking algorithms that can work without strong stochastic transitivity. We developed novel algorithms under this practical yet harder setting. Our efficient algorithm requires fewer human queries compared with algorithms designed with stronger assumptions. The algorithm is provably optimal. The second part focuses on preference learning without transitivity assumption. In reality, humans rarely make consistent comparisons and often demonstrate contradicting preferences such as a loop within the preference relations. I considered the most general setting where there is no transitivity at all. I proposed algorithms that identify the Borda winner, an optimal choice even when a true underlying rank does not exist. I showed the algorithm enjoys minimum regret, a notion that trades between exploration and exploitation. This result sheds light on the fundamental difficulty and cost of recovering human preferences under the fewest assumptions.The third part also focuses on preference learning without transitivity assumption, but instead considers an alternative definition of learning objective, the von Neumann winner. I first formulate the general preference as a game environment where two players aim to win over each other and then present an algorithmic framework that can solve this game in a self-play manner asymptotically. The algorithm is then extended to the task of fine-tuning large language models and shows remarkable empirical performance.The methods and techniques discussed in this dissertation cover a full spectrum of different assumptions and settings of preference learning. In each setting, the new algorithms are presented along with theoretical analysis ensuring a tight performance guarantee. Additionally, during the exploration of different settings, new research directions and open questions are identified, which could help promote the research of preference learning in terms of sample efficiency in the future.
- 일반주제명
- Computer science
- 일반주제명
- Statistics
- 일반주제명
- Information science
- 키워드
- Bandits
- 키워드
- Human feedback
- 키워드
- Ranking
- 기타저자
- University of California, Los Angeles Computer Science 0201
- 기본자료저록
- Dissertations Abstracts International. 86-04A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017165186
■00520250211153146
■006m o d
■007cr#unu||||||||
■020 ▼a9798384464778
■035 ▼a(MiAaPQ)AAI31563320
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aWu, Yue.
■24510▼aLearning From Human Feedback: Ranking, Bandit, and Preference Optimization
■260 ▼a[Sl]▼bUniversity of California, Los Angeles▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a188 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-04, Section: A.
■500 ▼aAdvisor: Gu, Quanquan.
■5021 ▼aThesis (Ph.D.)--University of California, Los Angeles, 2024.
■520 ▼aThis dissertation investigates several challenges in artificial intelligence (AI) alignment and reinforcement learning (RL), particularly focusing on applications when only preference feedback is available. Learning from preference feedback has been one central problem across different fields such as ranking, recommendation systems, and social choice theory. Recently, reinforcement learning from human feedback (RLHF) has also shown its strong potential in utilizing weakly supervised human data (preference feedback) and its ability to encode human values into machine learning models accurately. This dissertation aims to comprehensively characterize preference-based statistical learning, focusing on the sample complexity of ranking and preference model estimation and fine-tuning large language models.The first part of the dissertation explores novel methods for the learning-to-rank problem. I studied learning to rank under the strong stochastic transitivity (SST) condition, a prevalent model without assuming a score for each option. SST assumes that the accuracy of the comparison between two items increases as the disparity in their quality widens. I proposed one of the first adaptive approaches that can effectively aggregate the feedback from different human labelers and illustrated how the relationship between the number of human queries and resulting performance depends on the properties of the human labelers. I further follow up in this direction with my collaborators and provide active ranking algorithms that can work without strong stochastic transitivity. We developed novel algorithms under this practical yet harder setting. Our efficient algorithm requires fewer human queries compared with algorithms designed with stronger assumptions. The algorithm is provably optimal. The second part focuses on preference learning without transitivity assumption. In reality, humans rarely make consistent comparisons and often demonstrate contradicting preferences such as a loop within the preference relations. I considered the most general setting where there is no transitivity at all. I proposed algorithms that identify the Borda winner, an optimal choice even when a true underlying rank does not exist. I showed the algorithm enjoys minimum regret, a notion that trades between exploration and exploitation. This result sheds light on the fundamental difficulty and cost of recovering human preferences under the fewest assumptions.The third part also focuses on preference learning without transitivity assumption, but instead considers an alternative definition of learning objective, the von Neumann winner. I first formulate the general preference as a game environment where two players aim to win over each other and then present an algorithmic framework that can solve this game in a self-play manner asymptotically. The algorithm is then extended to the task of fine-tuning large language models and shows remarkable empirical performance.The methods and techniques discussed in this dissertation cover a full spectrum of different assumptions and settings of preference learning. In each setting, the new algorithms are presented along with theoretical analysis ensuring a tight performance guarantee. Additionally, during the exploration of different settings, new research directions and open questions are identified, which could help promote the research of preference learning in terms of sample efficiency in the future.
■590 ▼aSchool code: 0031.
■650 4▼aComputer science
■650 4▼aStatistics
■650 4▼aInformation science
■653 ▼aBandits
■653 ▼aHuman feedback
■653 ▼aRanking
■653 ▼aReinforcement learning
■653 ▼aPreference optimization
■690 ▼a0800
■690 ▼a0984
■690 ▼a0723
■690 ▼a0463
■71020▼aUniversity of California, Los Angeles▼bComputer Science 0201.
■7730 ▼tDissertations Abstracts International▼g86-04A.
■790 ▼a0031
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17165186▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


