서브메뉴
검색
Text Authorship in the Age of Large Language Models
Text Authorship in the Age of Large Language Models
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103045
- ISBN
- 9798311915182
- DDC
- 000
- 서명/저자
- Text Authorship in the Age of Large Language Models
- 발행사항
- [Sl] : The Pennsylvania State University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 156 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-11, Section: B.
- 주기사항
- Advisor: Lee, Dongwon.
- 학위논문주기
- Thesis (Ph.D.)--The Pennsylvania State University, 2024.
- 초록/해제
- 요약Tremendous progress in text generation by Large Language Models (LLMs) has led to an exponential rise in both the quality and quantity of LLM-generated texts. We are now surrounded by texts that are written entirely or enhanced and edited by autoregressive models. These texts appear in many contexts, ranging from dialog turns in an interactive session with ChatGPT to academic articles summarized by an LLM, a news article generated entirely by a model on social media, and so on. The ubiquity and high quality of such texts have made tracking and detecting their presence a task of growing and urgent importance. Particularly, there are developing concerns about copyright infringement, privacy, malicious use, intellectual property (IP) rights, and academic integrity that require active efforts to identify and trace LLM texts.In this thesis, we study machine-generated texts through four authorship-related tasks: (1) Human v/s Machine-generated Text evaluation: We first study if machines have human-text-like traits as measured by psycholinguistics-based measures. To do this, we turn to the Uniform Information Density (UID) principle that states that humans tend to distribute information or surprisal evenly or smoothly in language production. We analyze if machine-generated texts follow similar surprisal patterns and find that the answer depends on the decoding strategy used, with some settings generating more "human-like" surprisal distributions than others. But overall, we find that machines distribute surprisal differently than humans. Building upon this, we move to the next task, (2) Machine-generated text detection and Authorship Attribution: We develop "GPT-who", an authorship attributor that uses surprisal-based features to identify if a text is human-written or machine-generated, and also predicts the exact author LLM. We then study the reverse problem of (3) Authorship Obfuscation where the goal is to obfuscate or hide an author's identity by preserving semantics but altering the writing style such that it cannot be traced back to the original author. To do this, we present "ALISON", an obfuscation method that perturbs individual authors' syntactic patterns.Beyond single-authored texts, we explore (4) multi-LLM collaborative text generation by creating "CollabStory", a benchmark dataset containing over 35k creative stories generated jointly by up to 5 state-of-the-art open-source LLMs. We do this to study how authorship-related tasks evolve when multiple authors are present in a text, in light of unifying frameworks such as vLLM and LangChain that have enabled this oncoming scenario. Through these novel methods and datasets, this dissertation advances our understanding of authorship in the evolving landscape of LLM-generated texts and provides practical tools for addressing emerging challenges.
- 일반주제명
- Authorship
- 일반주제명
- Success
- 일반주제명
- Large language models
- 일반주제명
- Entropy
- 일반주제명
- Semantics
- 일반주제명
- Computer science
- 일반주제명
- Linguistics
- 일반주제명
- Information science
- 키워드
- Text authorship
- 기본자료저록
- Dissertations Abstracts International. 86-11B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2024 us c eng d■001000017356833
■00520260202103045
■006m o d
■007cr#unu||||||||
■020 ▼a9798311915182
■035 ▼a(MiAaPQ)AAI31897554
■035 ▼a(MiAaPQ)PennState22179szv4
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a000
■1001 ▼aVenkatraman, Saranya.
■24510▼aText Authorship in the Age of Large Language Models
■260 ▼a[Sl]▼bThe Pennsylvania State University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a156 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-11, Section: B.
■500 ▼aAdvisor: Lee, Dongwon.
■5021 ▼aThesis (Ph.D.)--The Pennsylvania State University, 2024.
■520 ▼aTremendous progress in text generation by Large Language Models (LLMs) has led to an exponential rise in both the quality and quantity of LLM-generated texts. We are now surrounded by texts that are written entirely or enhanced and edited by autoregressive models. These texts appear in many contexts, ranging from dialog turns in an interactive session with ChatGPT to academic articles summarized by an LLM, a news article generated entirely by a model on social media, and so on. The ubiquity and high quality of such texts have made tracking and detecting their presence a task of growing and urgent importance. Particularly, there are developing concerns about copyright infringement, privacy, malicious use, intellectual property (IP) rights, and academic integrity that require active efforts to identify and trace LLM texts.In this thesis, we study machine-generated texts through four authorship-related tasks: (1) Human v/s Machine-generated Text evaluation: We first study if machines have human-text-like traits as measured by psycholinguistics-based measures. To do this, we turn to the Uniform Information Density (UID) principle that states that humans tend to distribute information or surprisal evenly or smoothly in language production. We analyze if machine-generated texts follow similar surprisal patterns and find that the answer depends on the decoding strategy used, with some settings generating more "human-like" surprisal distributions than others. But overall, we find that machines distribute surprisal differently than humans. Building upon this, we move to the next task, (2) Machine-generated text detection and Authorship Attribution: We develop "GPT-who", an authorship attributor that uses surprisal-based features to identify if a text is human-written or machine-generated, and also predicts the exact author LLM. We then study the reverse problem of (3) Authorship Obfuscation where the goal is to obfuscate or hide an author's identity by preserving semantics but altering the writing style such that it cannot be traced back to the original author. To do this, we present "ALISON", an obfuscation method that perturbs individual authors' syntactic patterns.Beyond single-authored texts, we explore (4) multi-LLM collaborative text generation by creating "CollabStory", a benchmark dataset containing over 35k creative stories generated jointly by up to 5 state-of-the-art open-source LLMs. We do this to study how authorship-related tasks evolve when multiple authors are present in a text, in light of unifying frameworks such as vLLM and LangChain that have enabled this oncoming scenario. Through these novel methods and datasets, this dissertation advances our understanding of authorship in the evolving landscape of LLM-generated texts and provides practical tools for addressing emerging challenges.
■590 ▼aSchool code: 0176.
■650 4▼aAuthorship
■650 4▼aSuccess
■650 4▼aLarge language models
■650 4▼aEntropy
■650 4▼aSemantics
■650 4▼aComputer science
■650 4▼aLinguistics
■650 4▼aInformation science
■653 ▼aLarge Language Models
■653 ▼aUniform Information Density
■653 ▼aText authorship
■690 ▼a0290
■690 ▼a0984
■690 ▼a0723
■71020▼aThe Pennsylvania State University.
■7730 ▼tDissertations Abstracts International▼g86-11B.
■790 ▼a0176
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17356833▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


