본문

서브메뉴

The Moral Alignment of Large Language Models
The Moral Alignment of Large Language Models
The Moral Alignment of Large Language Models

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202104704
ISBN  
9798291555262
DDC  
301.1
저자명  
Dillion, Danica.
서명/저자  
The Moral Alignment of Large Language Models
발행사항  
[Sl] : The University of North Carolina at Chapel Hill, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
338 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-02, Section: B.
주기사항  
Advisor: Gray, Kurt.
학위논문주기  
Thesis (Ph.D.)--The University of North Carolina at Chapel Hill, 2025.
초록/해제  
요약Since the advent of machines, people have questioned whether they could ever model the complexities of human morality. New AI systems like large language models (LLMs) show promise in addressing this challenge, and their growing presence in everyday life makes the pursuit of moral alignment an urgent goal. This dissertation examines the extent to which LLMs align with human moral reasoning and explores strategies to enhance their moral alignment. Across three papers, I evaluate LLMs on three key dimensions: judgment alignment (how closely LLMs' moral judgments match those of people), explanation alignment (how clearly LLMs can explain their moral judgments), and cognitive alignment (how closely LLM "reasoning" resembles the processes people use to make moral judgments). Paper 1 compares LLM moral ratings on a diverse set of real-world scenarios to ratings provided by American participants, revealing high correlations between the two. Paper 2 assesses LLM explanation alignment by comparing perceptions of LLM-generated moral explanations and advice with those written by a representative sample of Americans and a professional ethicist. Participants preferred the LLMs' explanations and advice, rating them as more morally sound, trustworthy, thoughtful, and correct. Paper 3 introduces a method to enhance cognitive alignment through a "moral bottleneck" that prompts models to explicitly consider a moral template similar to that used by people before making judgments. I find that this method improves the perceived transparency of LLM moral decision-making while typically maintaining or enhancing their alignment with human moral judgments. Together, these studies suggest that LLMs can model moral judgments with relatively high fidelity, and that some of their remaining limitations can be addressed through psychologically informed interventions.
일반주제명  
Social psychology
일반주제명  
Neurosciences
일반주제명  
Psychology
키워드  
Alignment
키워드  
Large language models
키워드  
Morality
키워드  
Moral judgments
기타저자  
The University of North Carolina at Chapel Hill Psychology
기본자료저록  
Dissertations Abstracts International. 87-02B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017358454
■00520260202104704
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798291555262
■035    ▼a(MiAaPQ)AAI32116815
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a301.1
■1001  ▼aDillion,  Danica.
■24510▼aThe  Moral  Alignment  of  Large  Language  Models
■260    ▼a[Sl]▼bThe  University  of  North  Carolina  at  Chapel  Hill▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a338  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-02,  Section:  B.
■500    ▼aAdvisor:  Gray,  Kurt.
■5021  ▼aThesis  (Ph.D.)--The  University  of  North  Carolina  at  Chapel  Hill,  2025.
■520    ▼aSince  the  advent  of  machines,  people  have  questioned  whether  they  could  ever  model  the  complexities  of  human  morality.  New  AI  systems  like  large  language  models  (LLMs)  show  promise  in  addressing  this  challenge,  and  their  growing  presence  in  everyday  life  makes  the  pursuit  of  moral  alignment  an  urgent  goal.  This  dissertation  examines  the  extent  to  which  LLMs  align  with  human  moral  reasoning  and  explores  strategies  to  enhance  their  moral  alignment.  Across  three  papers,  I  evaluate  LLMs  on  three  key  dimensions:  judgment  alignment  (how  closely  LLMs'  moral  judgments  match  those  of  people),  explanation  alignment  (how  clearly  LLMs  can  explain  their  moral  judgments),  and  cognitive  alignment  (how  closely  LLM  "reasoning"  resembles  the  processes  people  use  to  make  moral  judgments).  Paper  1  compares  LLM  moral  ratings  on  a  diverse  set  of  real-world  scenarios  to  ratings  provided  by  American  participants,  revealing  high  correlations  between  the  two.  Paper  2  assesses  LLM  explanation  alignment  by  comparing  perceptions  of  LLM-generated  moral  explanations  and  advice  with  those  written  by  a  representative  sample  of  Americans  and  a  professional  ethicist.  Participants  preferred  the  LLMs'  explanations  and  advice,  rating  them  as  more  morally  sound,  trustworthy,  thoughtful,  and  correct.  Paper  3  introduces  a  method  to  enhance  cognitive  alignment  through  a  "moral  bottleneck"  that  prompts  models  to  explicitly  consider  a  moral  template  similar  to  that  used  by  people  before  making  judgments.  I  find  that  this  method  improves  the  perceived  transparency  of  LLM  moral  decision-making  while  typically  maintaining  or  enhancing  their  alignment  with  human  moral  judgments.  Together,  these  studies  suggest  that  LLMs  can  model  moral  judgments  with  relatively  high  fidelity,  and  that  some  of  their  remaining  limitations  can  be  addressed  through  psychologically  informed  interventions.
■590    ▼aSchool  code:  0153.
■650  4▼aSocial  psychology
■650  4▼aNeurosciences
■650  4▼aPsychology
■653    ▼aAlignment
■653    ▼aLarge  language  models
■653    ▼aMorality
■653    ▼aMoral  judgments
■690    ▼a0451
■690    ▼a0800
■690    ▼a0621
■690    ▼a0317
■71020▼aThe  University  of  North  Carolina  at  Chapel  Hill▼bPsychology.
■7730  ▼tDissertations  Abstracts  International▼g87-02B.
■790    ▼a0153
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17358454▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF16032 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.