서브메뉴
검색
Scalable Alignment of Large Language Models Towards Truth Seeking, Complex Reasoning, and Human Values
Scalable Alignment of Large Language Models Towards Truth Seeking, Complex Reasoning, and Human Values
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103150
- ISBN
- 9798314864319
- DDC
- 004
- 저자명
- Sun, Zhiqing.
- 서명/저자
- Scalable Alignment of Large Language Models Towards Truth Seeking, Complex Reasoning, and Human Values
- 발행사항
- [Sl] : Carnegie Mellon University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 147 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-11, Section: B.
- 주기사항
- Advisor: Yang, Yiming.
- 학위논문주기
- Thesis (Ph.D.)--Carnegie Mellon University, 2025.
- 초록/해제
- 요약The exponential advancement in Large Language Models (LLMs) and reasoning-powered AI agents, exemplified by GPT-4 and OpenAI Deep Research, has accelerated the timeline toward Artificial General Intelligence (AGI), with capabilities expanding at an unprecedented rate. As we stand at the threshold of potentially achieving AGI in the near future, the challenge of alignment-ensuring these systems remain truthful, capable of sophisticated reasoning, and aligned with human values-has become increasingly critical.This thesis proposes novel methodologies to address fundamental alignment challenges for systems approaching superhuman capabilities. Extending beyond conventional paradigms such as Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), we develop scalable alignment mechanisms through our Principle-Driven Alignment methodology. Implemented within a reinforcement learning from AI feedback (RLAIF) framework, this approach demonstrates significant improvements in maintaining system reliability under capability scaling. To mitigate factual inconsistencies in generation, we introduce Recitation Augmentation and Factually Augmented RLHF, which demonstrate robust performance on large language and multimodal models. The proposed Easy-to-Hard Generalization framework provides a systematic approach for preserving alignment by leveraging the insight that models can more reliably evaluate solutions than generate them, enabling supervision of complex reasoning tasks through reward models trained on simpler problems. Additionally, we proposed Lean-STaR, a framework that improves theorem-proving performance by guiding models to generate informal thoughts before formal solutions, demonstrating the effectiveness of Chain-of-Thought reasoning in enhancing autonomous decision-making capabilities while providing greater transparency of model reasoning processes.This research contributes to a critical area of AI development by establishing rigorous frameworks for maintaining alignment as systems become increasingly capable. Our findings demonstrate the effectiveness of these approaches in creating AI systems that are aligned with fundamental human values while preserving performance reliability. These frameworks provide a foundation for scalable solutions that will shape the future development of advanced AI systems.
- 일반주제명
- Computer science
- 일반주제명
- Information technology
- 키워드
- AI alignment
- 키워드
- AI reasoning
- 기타저자
- Carnegie Mellon University Language Technologies Institute
- 기본자료저록
- Dissertations Abstracts International. 86-11B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017357218
■00520260202103150
■006m o d
■007cr#unu||||||||
■020 ▼a9798314864319
■035 ▼a(MiAaPQ)AAI31996440
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aSun, Zhiqing.
■24510▼aScalable Alignment of Large Language Models Towards Truth Seeking, Complex Reasoning, and Human Values
■260 ▼a[Sl]▼bCarnegie Mellon University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a147 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-11, Section: B.
■500 ▼aAdvisor: Yang, Yiming.
■5021 ▼aThesis (Ph.D.)--Carnegie Mellon University, 2025.
■520 ▼aThe exponential advancement in Large Language Models (LLMs) and reasoning-powered AI agents, exemplified by GPT-4 and OpenAI Deep Research, has accelerated the timeline toward Artificial General Intelligence (AGI), with capabilities expanding at an unprecedented rate. As we stand at the threshold of potentially achieving AGI in the near future, the challenge of alignment-ensuring these systems remain truthful, capable of sophisticated reasoning, and aligned with human values-has become increasingly critical.This thesis proposes novel methodologies to address fundamental alignment challenges for systems approaching superhuman capabilities. Extending beyond conventional paradigms such as Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), we develop scalable alignment mechanisms through our Principle-Driven Alignment methodology. Implemented within a reinforcement learning from AI feedback (RLAIF) framework, this approach demonstrates significant improvements in maintaining system reliability under capability scaling. To mitigate factual inconsistencies in generation, we introduce Recitation Augmentation and Factually Augmented RLHF, which demonstrate robust performance on large language and multimodal models. The proposed Easy-to-Hard Generalization framework provides a systematic approach for preserving alignment by leveraging the insight that models can more reliably evaluate solutions than generate them, enabling supervision of complex reasoning tasks through reward models trained on simpler problems. Additionally, we proposed Lean-STaR, a framework that improves theorem-proving performance by guiding models to generate informal thoughts before formal solutions, demonstrating the effectiveness of Chain-of-Thought reasoning in enhancing autonomous decision-making capabilities while providing greater transparency of model reasoning processes.This research contributes to a critical area of AI development by establishing rigorous frameworks for maintaining alignment as systems become increasingly capable. Our findings demonstrate the effectiveness of these approaches in creating AI systems that are aligned with fundamental human values while preserving performance reliability. These frameworks provide a foundation for scalable solutions that will shape the future development of advanced AI systems.
■590 ▼aSchool code: 0041.
■650 4▼aComputer science
■650 4▼aInformation technology
■653 ▼aAI alignment
■653 ▼aAI reasoning
■653 ▼aLarge Language Models
■653 ▼aScalable oversight
■653 ▼aArtificial General Intelligence
■690 ▼a0800
■690 ▼a0984
■690 ▼a0489
■71020▼aCarnegie Mellon University▼bLanguage Technologies Institute.
■7730 ▼tDissertations Abstracts International▼g86-11B.
■790 ▼a0041
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357218▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


