본문

서브메뉴

Scalable Alignment of Large Language Models Towards Truth Seeking, Complex Reasoning, and Human Values
Scalable Alignment of Large Language Models Towards Truth Seeking, Complex Reasoning, and ...
Scalable Alignment of Large Language Models Towards Truth Seeking, Complex Reasoning, and Human Values

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202103150
ISBN  
9798314864319
DDC  
004
저자명  
Sun, Zhiqing.
서명/저자  
Scalable Alignment of Large Language Models Towards Truth Seeking, Complex Reasoning, and Human Values
발행사항  
[Sl] : Carnegie Mellon University, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
147 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-11, Section: B.
주기사항  
Advisor: Yang, Yiming.
학위논문주기  
Thesis (Ph.D.)--Carnegie Mellon University, 2025.
초록/해제  
요약The exponential advancement in Large Language Models (LLMs) and reasoning-powered AI agents, exemplified by GPT-4 and OpenAI Deep Research, has accelerated the timeline toward Artificial General Intelligence (AGI), with capabilities expanding at an unprecedented rate. As we stand at the threshold of potentially achieving AGI in the near future, the challenge of alignment-ensuring these systems remain truthful, capable of sophisticated reasoning, and aligned with human values-has become increasingly critical.This thesis proposes novel methodologies to address fundamental alignment challenges for systems approaching superhuman capabilities. Extending beyond conventional paradigms such as Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), we develop scalable alignment mechanisms through our Principle-Driven Alignment methodology. Implemented within a reinforcement learning from AI feedback (RLAIF) framework, this approach demonstrates significant improvements in maintaining system reliability under capability scaling. To mitigate factual inconsistencies in generation, we introduce Recitation Augmentation and Factually Augmented RLHF, which demonstrate robust performance on large language and multimodal models. The proposed Easy-to-Hard Generalization framework provides a systematic approach for preserving alignment by leveraging the insight that models can more reliably evaluate solutions than generate them, enabling supervision of complex reasoning tasks through reward models trained on simpler problems. Additionally, we proposed Lean-STaR, a framework that improves theorem-proving performance by guiding models to generate informal thoughts before formal solutions, demonstrating the effectiveness of Chain-of-Thought reasoning in enhancing autonomous decision-making capabilities while providing greater transparency of model reasoning processes.This research contributes to a critical area of AI development by establishing rigorous frameworks for maintaining alignment as systems become increasingly capable. Our findings demonstrate the effectiveness of these approaches in creating AI systems that are aligned with fundamental human values while preserving performance reliability. These frameworks provide a foundation for scalable solutions that will shape the future development of advanced AI systems.
일반주제명  
Computer science
일반주제명  
Information technology
키워드  
AI alignment
키워드  
AI reasoning
키워드  
Large Language Models
키워드  
Scalable oversight
키워드  
Artificial General Intelligence
기타저자  
Carnegie Mellon University Language Technologies Institute
기본자료저록  
Dissertations Abstracts International. 86-11B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017357218
■00520260202103150
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798314864319
■035    ▼a(MiAaPQ)AAI31996440
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aSun,  Zhiqing.
■24510▼aScalable  Alignment  of  Large  Language  Models  Towards  Truth  Seeking,  Complex  Reasoning,  and  Human  Values
■260    ▼a[Sl]▼bCarnegie  Mellon  University▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a147  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-11,  Section:  B.
■500    ▼aAdvisor:  Yang,  Yiming.
■5021  ▼aThesis  (Ph.D.)--Carnegie  Mellon  University,  2025.
■520    ▼aThe  exponential  advancement  in  Large  Language  Models  (LLMs)  and  reasoning-powered  AI  agents,  exemplified  by  GPT-4  and  OpenAI  Deep  Research,  has  accelerated  the  timeline  toward  Artificial  General  Intelligence  (AGI),  with  capabilities  expanding  at  an  unprecedented  rate.  As  we  stand  at  the  threshold  of  potentially  achieving  AGI  in  the  near  future,  the  challenge  of  alignment-ensuring  these  systems  remain  truthful,  capable  of  sophisticated  reasoning,  and  aligned  with  human  values-has  become  increasingly  critical.This  thesis  proposes  novel  methodologies  to  address  fundamental  alignment  challenges  for  systems  approaching  superhuman  capabilities.  Extending  beyond  conventional  paradigms  such  as  Supervised  Fine-Tuning  (SFT)  and  Reinforcement  Learning  from  Human  Feedback  (RLHF),  we  develop  scalable  alignment  mechanisms  through  our  Principle-Driven  Alignment  methodology.  Implemented  within  a  reinforcement  learning  from  AI  feedback  (RLAIF)  framework,  this  approach  demonstrates  significant  improvements  in  maintaining  system  reliability  under  capability  scaling.  To  mitigate  factual  inconsistencies  in  generation,  we  introduce  Recitation  Augmentation  and  Factually  Augmented  RLHF,  which  demonstrate  robust  performance  on  large  language  and  multimodal  models.  The  proposed  Easy-to-Hard  Generalization  framework  provides  a  systematic  approach  for  preserving  alignment  by  leveraging  the  insight  that  models  can  more  reliably  evaluate  solutions  than  generate  them,  enabling  supervision  of  complex  reasoning  tasks  through  reward  models  trained  on  simpler  problems.  Additionally,  we  proposed  Lean-STaR,  a  framework  that  improves  theorem-proving  performance  by  guiding  models  to  generate  informal  thoughts  before  formal  solutions,  demonstrating  the  effectiveness  of  Chain-of-Thought  reasoning  in  enhancing  autonomous  decision-making  capabilities  while  providing  greater  transparency  of  model  reasoning  processes.This  research  contributes  to  a  critical  area  of  AI  development  by  establishing  rigorous  frameworks  for  maintaining  alignment  as  systems  become  increasingly  capable.  Our  findings  demonstrate  the  effectiveness  of  these  approaches  in  creating  AI  systems  that  are  aligned  with  fundamental  human  values  while  preserving  performance  reliability.  These  frameworks  provide  a  foundation  for  scalable  solutions  that  will  shape  the  future  development  of  advanced  AI  systems.
■590    ▼aSchool  code:  0041.
■650  4▼aComputer  science
■650  4▼aInformation  technology
■653    ▼aAI  alignment
■653    ▼aAI  reasoning
■653    ▼aLarge  Language  Models
■653    ▼aScalable  oversight
■653    ▼aArtificial  General  Intelligence
■690    ▼a0800
■690    ▼a0984
■690    ▼a0489
■71020▼aCarnegie  Mellon  University▼bLanguage  Technologies  Institute.
■7730  ▼tDissertations  Abstracts  International▼g86-11B.
■790    ▼a0041
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357218▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF19059 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.