서브메뉴
검색
Defending Against Authorship Attribution Attacks With Large Language Models
Defending Against Authorship Attribution Attacks With Large Language Models
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202104654
- ISBN
- 9798286428632
- DDC
- 020
- 저자명
- Wang, Haining.
- 서명/저자
- Defending Against Authorship Attribution Attacks With Large Language Models
- 발행사항
- [Sl] : Indiana University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 236 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-12, Section: B.
- 주기사항
- Advisor: Riddell, Allen B.
- 학위논문주기
- Thesis (Ph.D.)--Indiana University, 2025.
- 초록/해제
- 요약In today's digital era, individuals leave significant digital footprints through their writing, whether on social media or on their employer's devices. These digital footprints pose a serious challenge for identity protection: authorship attribution techniques can identify the author of an unsigned document with high accuracy. This threat is especially acute for those who must speak publicly while safeguarding their anonymity, including whistleblowers, journalists, activists, and individuals living under oppressive regimes.Defenses against authorship attribution attacks rely on altering an individual's writing style, making it unlinkable to their prior work while maintaining meaning and fluency. Despite extensive efforts at automation, existing techniques rarely match the effectiveness of manual interventions and make significant technical demands of individuals seeking to obfuscate their writing style.This dissertation investigates the use of large language models (LLMs) as an effective defense against authorship attribution attacks. These models are user-friendly and respond directly to natural language prompts, making them particularly accessible for privacy-conscious individuals. Through extensive experiments, this dissertation reproduces both established automated and manual circumvention strategies with LLMs.The results confirm that, with the right prompts, LLMs can offer significant protection from authorship attribution attacks. A simple "write differently" prompt on lightweight LLMs produces semantically faithful, inconspicuous text while driving attribution models' performance down to near-chance levels. Surprisingly, open-weights models with just 8-9 billion parameters consistently outperform far larger closed-source models. Furthermore, this research overturns assumptions about in-context learning, showing that adding context, such as personas, exemplars, or extended demonstrations, often harms rather than helps defensive performance.These findings advance our understanding of how LLMs can frustrate stylometric fingerprinting while providing actionable guidance for those who need anonymization most, yet may struggle to access its benefits. At the same time, by bridging theory and practice, this dissertation delivers a practical solution to defend against authorship attribution attacks.
- 일반주제명
- Information science
- 일반주제명
- Library science
- 키워드
- Machine learning
- 키워드
- Writing style
- 기타저자
- Indiana University Information Science
- 기본자료저록
- Dissertations Abstracts International. 86-12B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017358388
■00520260202104654
■006m o d
■007cr#unu||||||||
■020 ▼a9798286428632
■035 ▼a(MiAaPQ)AAI32115646
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a020
■1001 ▼aWang, Haining.▼0(orcid)0000-0002-1196-0918
■24510▼aDefending Against Authorship Attribution Attacks With Large Language Models
■260 ▼a[Sl]▼bIndiana University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a236 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-12, Section: B.
■500 ▼aAdvisor: Riddell, Allen B.
■5021 ▼aThesis (Ph.D.)--Indiana University, 2025.
■520 ▼aIn today's digital era, individuals leave significant digital footprints through their writing, whether on social media or on their employer's devices. These digital footprints pose a serious challenge for identity protection: authorship attribution techniques can identify the author of an unsigned document with high accuracy. This threat is especially acute for those who must speak publicly while safeguarding their anonymity, including whistleblowers, journalists, activists, and individuals living under oppressive regimes.Defenses against authorship attribution attacks rely on altering an individual's writing style, making it unlinkable to their prior work while maintaining meaning and fluency. Despite extensive efforts at automation, existing techniques rarely match the effectiveness of manual interventions and make significant technical demands of individuals seeking to obfuscate their writing style.This dissertation investigates the use of large language models (LLMs) as an effective defense against authorship attribution attacks. These models are user-friendly and respond directly to natural language prompts, making them particularly accessible for privacy-conscious individuals. Through extensive experiments, this dissertation reproduces both established automated and manual circumvention strategies with LLMs.The results confirm that, with the right prompts, LLMs can offer significant protection from authorship attribution attacks. A simple "write differently" prompt on lightweight LLMs produces semantically faithful, inconspicuous text while driving attribution models' performance down to near-chance levels. Surprisingly, open-weights models with just 8-9 billion parameters consistently outperform far larger closed-source models. Furthermore, this research overturns assumptions about in-context learning, showing that adding context, such as personas, exemplars, or extended demonstrations, often harms rather than helps defensive performance.These findings advance our understanding of how LLMs can frustrate stylometric fingerprinting while providing actionable guidance for those who need anonymization most, yet may struggle to access its benefits. At the same time, by bridging theory and practice, this dissertation delivers a practical solution to defend against authorship attribution attacks.
■590 ▼aSchool code: 0093.
■650 4▼aInformation science
■650 4▼aLibrary science
■653 ▼aAdversarial stylometry
■653 ▼aAuthorship attribution
■653 ▼aMachine learning
■653 ▼aNatural language processing
■653 ▼aWriting style
■690 ▼a0723
■690 ▼a0399
■690 ▼a0800
■71020▼aIndiana University▼bInformation Science.
■7730 ▼tDissertations Abstracts International▼g86-12B.
■790 ▼a0093
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17358388▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


