서브메뉴
검색
The Moral Alignment of Large Language Models
The Moral Alignment of Large Language Models
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202104704
- ISBN
- 9798291555262
- DDC
- 301.1
- 저자명
- Dillion, Danica.
- 서명/저자
- The Moral Alignment of Large Language Models
- 발행사항
- [Sl] : The University of North Carolina at Chapel Hill, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 338 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-02, Section: B.
- 주기사항
- Advisor: Gray, Kurt.
- 학위논문주기
- Thesis (Ph.D.)--The University of North Carolina at Chapel Hill, 2025.
- 초록/해제
- 요약Since the advent of machines, people have questioned whether they could ever model the complexities of human morality. New AI systems like large language models (LLMs) show promise in addressing this challenge, and their growing presence in everyday life makes the pursuit of moral alignment an urgent goal. This dissertation examines the extent to which LLMs align with human moral reasoning and explores strategies to enhance their moral alignment. Across three papers, I evaluate LLMs on three key dimensions: judgment alignment (how closely LLMs' moral judgments match those of people), explanation alignment (how clearly LLMs can explain their moral judgments), and cognitive alignment (how closely LLM "reasoning" resembles the processes people use to make moral judgments). Paper 1 compares LLM moral ratings on a diverse set of real-world scenarios to ratings provided by American participants, revealing high correlations between the two. Paper 2 assesses LLM explanation alignment by comparing perceptions of LLM-generated moral explanations and advice with those written by a representative sample of Americans and a professional ethicist. Participants preferred the LLMs' explanations and advice, rating them as more morally sound, trustworthy, thoughtful, and correct. Paper 3 introduces a method to enhance cognitive alignment through a "moral bottleneck" that prompts models to explicitly consider a moral template similar to that used by people before making judgments. I find that this method improves the perceived transparency of LLM moral decision-making while typically maintaining or enhancing their alignment with human moral judgments. Together, these studies suggest that LLMs can model moral judgments with relatively high fidelity, and that some of their remaining limitations can be addressed through psychologically informed interventions.
- 일반주제명
- Social psychology
- 일반주제명
- Neurosciences
- 일반주제명
- Psychology
- 키워드
- Alignment
- 키워드
- Morality
- 키워드
- Moral judgments
- 기타저자
- The University of North Carolina at Chapel Hill Psychology
- 기본자료저록
- Dissertations Abstracts International. 87-02B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017358454
■00520260202104704
■006m o d
■007cr#unu||||||||
■020 ▼a9798291555262
■035 ▼a(MiAaPQ)AAI32116815
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a301.1
■1001 ▼aDillion, Danica.
■24510▼aThe Moral Alignment of Large Language Models
■260 ▼a[Sl]▼bThe University of North Carolina at Chapel Hill▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a338 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-02, Section: B.
■500 ▼aAdvisor: Gray, Kurt.
■5021 ▼aThesis (Ph.D.)--The University of North Carolina at Chapel Hill, 2025.
■520 ▼aSince the advent of machines, people have questioned whether they could ever model the complexities of human morality. New AI systems like large language models (LLMs) show promise in addressing this challenge, and their growing presence in everyday life makes the pursuit of moral alignment an urgent goal. This dissertation examines the extent to which LLMs align with human moral reasoning and explores strategies to enhance their moral alignment. Across three papers, I evaluate LLMs on three key dimensions: judgment alignment (how closely LLMs' moral judgments match those of people), explanation alignment (how clearly LLMs can explain their moral judgments), and cognitive alignment (how closely LLM "reasoning" resembles the processes people use to make moral judgments). Paper 1 compares LLM moral ratings on a diverse set of real-world scenarios to ratings provided by American participants, revealing high correlations between the two. Paper 2 assesses LLM explanation alignment by comparing perceptions of LLM-generated moral explanations and advice with those written by a representative sample of Americans and a professional ethicist. Participants preferred the LLMs' explanations and advice, rating them as more morally sound, trustworthy, thoughtful, and correct. Paper 3 introduces a method to enhance cognitive alignment through a "moral bottleneck" that prompts models to explicitly consider a moral template similar to that used by people before making judgments. I find that this method improves the perceived transparency of LLM moral decision-making while typically maintaining or enhancing their alignment with human moral judgments. Together, these studies suggest that LLMs can model moral judgments with relatively high fidelity, and that some of their remaining limitations can be addressed through psychologically informed interventions.
■590 ▼aSchool code: 0153.
■650 4▼aSocial psychology
■650 4▼aNeurosciences
■650 4▼aPsychology
■653 ▼aAlignment
■653 ▼aLarge language models
■653 ▼aMorality
■653 ▼aMoral judgments
■690 ▼a0451
■690 ▼a0800
■690 ▼a0621
■690 ▼a0317
■71020▼aThe University of North Carolina at Chapel Hill▼bPsychology.
■7730 ▼tDissertations Abstracts International▼g87-02B.
■790 ▼a0153
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17358454▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


