서브메뉴
검색
Human-Centered AI in Computational Social Science: Evaluating Automated Annotation with Large Language Models
Human-Centered AI in Computational Social Science: Evaluating Automated Annotation with Large Language Models
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103003
- ISBN
- 9798280760547
- DDC
- 320
- 서명/저자
- Human-Centered AI in Computational Social Science: Evaluating Automated Annotation with Large Language Models
- 발행사항
- [Sl] : University of Pennsylvania, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 147 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-12, Section: B.
- 주기사항
- Advisor: Hopkins, Daniel.
- 학위논문주기
- Thesis (Ph.D.)--University of Pennsylvania, 2025.
- 초록/해제
- 요약Computational social scientists are increasingly incorporating text as data into their research. A typical framework for working with large text data sets involves hiring human annotators to read a subset of the text samples and then building a statistical model to annotate the remainder of the text corpus. Due to their effectiveness at quantifying natural language, their ease of application, and their relatively low cost, artificial intelligence tools, like generative large language models (LLMs), may be used to automate these manual annotation procedures. This process, which I call "automated annotation," can dramatically improve research designs that involve text as data. For example, I demonstrate that automated annotation procedures can cost 11.6% that of standard annotation approaches and take 18.8% the time. Although automated annotation has remarkable potential in social science, there are serious concerns about misuse and uncritical application. If practitioners use automated annotation without validation, for instance, they risk unknown bias and other inaccuracies in downstream applications. Thus, my dissertation aims to test strategies to develop effective and responsible automated annotation procedures. Specifically, I argue for a human-centered automated annotation framework, which places a central role for human annotations at each stage of the workflow. Across three studies, I develop and implement various automated annotation techniques that all remain grounded in human reasoning. My empirical investigations cover a wide range of topics-from testing automated annotation strategies with generative LLMs to developing a multi-stage, human-in-the-loop annotation pipeline. As a whole, my findings underscore the potential of leveraging AI tools to enhance text-as-data methodologies and to help researchers explore important substantive questions. With proper validation techniques, generative LLMs can approximate human reasoning at a rapid pace and low cost.
- 일반주제명
- Political science
- 일반주제명
- Computer science
- 키워드
- Human reasoning
- 기타저자
- University of Pennsylvania Political Science
- 기본자료저록
- Dissertations Abstracts International. 86-12B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017356616
■00520260202103003
■006m o d
■007cr#unu||||||||
■020 ▼a9798280760547
■035 ▼a(MiAaPQ)AAI31840539
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a320
■1001 ▼aPangakis, Nicholas James.
■24510▼aHuman-Centered AI in Computational Social Science: Evaluating Automated Annotation with Large Language Models
■260 ▼a[Sl]▼bUniversity of Pennsylvania▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a147 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-12, Section: B.
■500 ▼aAdvisor: Hopkins, Daniel.
■5021 ▼aThesis (Ph.D.)--University of Pennsylvania, 2025.
■520 ▼aComputational social scientists are increasingly incorporating text as data into their research. A typical framework for working with large text data sets involves hiring human annotators to read a subset of the text samples and then building a statistical model to annotate the remainder of the text corpus. Due to their effectiveness at quantifying natural language, their ease of application, and their relatively low cost, artificial intelligence tools, like generative large language models (LLMs), may be used to automate these manual annotation procedures. This process, which I call "automated annotation," can dramatically improve research designs that involve text as data. For example, I demonstrate that automated annotation procedures can cost 11.6% that of standard annotation approaches and take 18.8% the time. Although automated annotation has remarkable potential in social science, there are serious concerns about misuse and uncritical application. If practitioners use automated annotation without validation, for instance, they risk unknown bias and other inaccuracies in downstream applications. Thus, my dissertation aims to test strategies to develop effective and responsible automated annotation procedures. Specifically, I argue for a human-centered automated annotation framework, which places a central role for human annotations at each stage of the workflow. Across three studies, I develop and implement various automated annotation techniques that all remain grounded in human reasoning. My empirical investigations cover a wide range of topics-from testing automated annotation strategies with generative LLMs to developing a multi-stage, human-in-the-loop annotation pipeline. As a whole, my findings underscore the potential of leveraging AI tools to enhance text-as-data methodologies and to help researchers explore important substantive questions. With proper validation techniques, generative LLMs can approximate human reasoning at a rapid pace and low cost.
■590 ▼aSchool code: 0175.
■650 4▼aPolitical science
■650 4▼aComputer science
■653 ▼aArtificial intelligence
■653 ▼aAutomated annotation
■653 ▼aComputational social science
■653 ▼aLarge language models
■653 ▼aHuman reasoning
■690 ▼a0615
■690 ▼a0984
■690 ▼a0800
■71020▼aUniversity of Pennsylvania▼bPolitical Science.
■7730 ▼tDissertations Abstracts International▼g86-12B.
■790 ▼a0175
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17356616▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


