본문

서브메뉴

Text Authorship in the Age of Large Language Models
Text Authorship in the Age of Large Language Models
Text Authorship in the Age of Large Language Models

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202103045
ISBN  
9798311915182
DDC  
000
저자명  
Venkatraman, Saranya.
서명/저자  
Text Authorship in the Age of Large Language Models
발행사항  
[Sl] : The Pennsylvania State University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
156 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-11, Section: B.
주기사항  
Advisor: Lee, Dongwon.
학위논문주기  
Thesis (Ph.D.)--The Pennsylvania State University, 2024.
초록/해제  
요약Tremendous progress in text generation by Large Language Models (LLMs) has led to an exponential rise in both the quality and quantity of LLM-generated texts. We are now surrounded by texts that are written entirely or enhanced and edited by autoregressive models. These texts appear in many contexts, ranging from dialog turns in an interactive session with ChatGPT to academic articles summarized by an LLM, a news article generated entirely by a model on social media, and so on. The ubiquity and high quality of such texts have made tracking and detecting their presence a task of growing and urgent importance. Particularly, there are developing concerns about copyright infringement, privacy, malicious use, intellectual property (IP) rights, and academic integrity that require active efforts to identify and trace LLM texts.In this thesis, we study machine-generated texts through four authorship-related tasks: (1) Human v/s Machine-generated Text evaluation: We first study if machines have human-text-like traits as measured by psycholinguistics-based measures. To do this, we turn to the Uniform Information Density (UID) principle that states that humans tend to distribute information or surprisal evenly or smoothly in language production. We analyze if machine-generated texts follow similar surprisal patterns and find that the answer depends on the decoding strategy used, with some settings generating more "human-like" surprisal distributions than others. But overall, we find that machines distribute surprisal differently than humans. Building upon this, we move to the next task, (2) Machine-generated text detection and Authorship Attribution: We develop "GPT-who", an authorship attributor that uses surprisal-based features to identify if a text is human-written or machine-generated, and also predicts the exact author LLM. We then study the reverse problem of (3) Authorship Obfuscation where the goal is to obfuscate or hide an author's identity by preserving semantics but altering the writing style such that it cannot be traced back to the original author. To do this, we present "ALISON", an obfuscation method that perturbs individual authors' syntactic patterns.Beyond single-authored texts, we explore (4) multi-LLM collaborative text generation by creating "CollabStory", a benchmark dataset containing over 35k creative stories generated jointly by up to 5 state-of-the-art open-source LLMs. We do this to study how authorship-related tasks evolve when multiple authors are present in a text, in light of unifying frameworks such as vLLM and LangChain that have enabled this oncoming scenario. Through these novel methods and datasets, this dissertation advances our understanding of authorship in the evolving landscape of LLM-generated texts and provides practical tools for addressing emerging challenges.
일반주제명  
Authorship
일반주제명  
Success
일반주제명  
Large language models
일반주제명  
Entropy
일반주제명  
Semantics
일반주제명  
Computer science
일반주제명  
Linguistics
일반주제명  
Information science
키워드  
Large Language Models
키워드  
Uniform Information Density
키워드  
Text authorship
기타저자  
The Pennsylvania State University.
기본자료저록  
Dissertations Abstracts International. 86-11B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2024        us                              c    eng  d
■001000017356833
■00520260202103045
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798311915182
■035    ▼a(MiAaPQ)AAI31897554
■035    ▼a(MiAaPQ)PennState22179szv4
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a000
■1001  ▼aVenkatraman,  Saranya.
■24510▼aText  Authorship  in  the  Age  of  Large  Language  Models
■260    ▼a[Sl]▼bThe  Pennsylvania  State  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a156  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-11,  Section:  B.
■500    ▼aAdvisor:  Lee,  Dongwon.
■5021  ▼aThesis  (Ph.D.)--The  Pennsylvania  State  University,  2024.
■520    ▼aTremendous  progress  in  text  generation  by  Large  Language  Models  (LLMs)  has  led  to  an  exponential  rise  in  both  the  quality  and  quantity  of  LLM-generated  texts.  We  are  now  surrounded  by  texts  that  are  written  entirely  or  enhanced  and  edited  by  autoregressive  models.  These  texts  appear  in  many  contexts,  ranging  from  dialog  turns  in  an  interactive  session  with  ChatGPT  to  academic  articles  summarized  by  an  LLM,  a  news  article  generated  entirely  by  a  model  on  social  media,  and  so  on.  The  ubiquity  and  high  quality  of  such  texts  have  made  tracking  and  detecting  their  presence  a  task  of  growing  and  urgent  importance.  Particularly,  there  are  developing  concerns  about  copyright  infringement,  privacy,  malicious  use,  intellectual  property  (IP)  rights,  and  academic  integrity  that  require  active  efforts  to  identify  and  trace  LLM  texts.In  this  thesis,  we  study  machine-generated  texts  through  four  authorship-related  tasks:  (1)  Human  v/s  Machine-generated  Text  evaluation:  We  first  study  if  machines  have  human-text-like  traits  as  measured  by  psycholinguistics-based  measures.  To  do  this,  we  turn  to  the  Uniform  Information  Density  (UID)  principle  that  states  that  humans  tend  to  distribute  information  or  surprisal  evenly  or  smoothly  in  language  production.  We  analyze  if  machine-generated  texts  follow  similar  surprisal  patterns  and  find  that  the  answer  depends  on  the  decoding  strategy  used,  with  some  settings  generating  more  "human-like"  surprisal  distributions  than  others.  But  overall,  we  find  that  machines  distribute  surprisal  differently  than  humans.  Building  upon  this,  we  move  to  the  next  task,  (2)  Machine-generated  text  detection  and  Authorship  Attribution:  We  develop  "GPT-who",  an  authorship  attributor  that  uses  surprisal-based  features  to  identify  if  a  text  is  human-written  or  machine-generated,  and  also  predicts  the  exact  author  LLM.  We  then  study  the  reverse  problem  of  (3)  Authorship  Obfuscation  where  the  goal  is  to  obfuscate  or  hide  an  author's  identity  by  preserving  semantics  but  altering  the  writing  style  such  that  it  cannot  be  traced  back  to  the  original  author.  To  do  this,  we  present  "ALISON",  an  obfuscation  method  that  perturbs  individual  authors'  syntactic  patterns.Beyond  single-authored  texts,  we  explore  (4)  multi-LLM  collaborative  text  generation  by  creating  "CollabStory",  a  benchmark  dataset  containing  over  35k  creative  stories  generated  jointly  by  up  to  5  state-of-the-art  open-source  LLMs.  We  do  this  to  study  how  authorship-related  tasks  evolve  when  multiple  authors  are  present  in  a  text,  in  light  of  unifying  frameworks  such  as  vLLM  and  LangChain  that  have  enabled  this  oncoming  scenario.  Through  these  novel  methods  and  datasets,  this  dissertation  advances  our  understanding  of  authorship  in  the  evolving  landscape  of  LLM-generated  texts  and  provides  practical  tools  for  addressing  emerging  challenges.
■590    ▼aSchool  code:  0176.
■650  4▼aAuthorship
■650  4▼aSuccess
■650  4▼aLarge  language  models
■650  4▼aEntropy
■650  4▼aSemantics
■650  4▼aComputer  science
■650  4▼aLinguistics
■650  4▼aInformation  science
■653    ▼aLarge  Language  Models
■653    ▼aUniform  Information  Density
■653    ▼aText  authorship
■690    ▼a0290
■690    ▼a0984
■690    ▼a0723
■71020▼aThe  Pennsylvania  State  University.
■7730  ▼tDissertations  Abstracts  International▼g86-11B.
■790    ▼a0176
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17356833▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF18492 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.