서브메뉴
검색
Large Language Models for Automatic Peer Review and Revision in Scientific Documents
Large Language Models for Automatic Peer Review and Revision in Scientific Documents
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211150920
- ISBN
- 9798381975161
- DDC
- 004
- 저자명
- D'Arcy, Mike.
- 서명/저자
- Large Language Models for Automatic Peer Review and Revision in Scientific Documents
- 발행사항
- [Sl] : Northwestern University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 170 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 85-10, Section: B.
- 주기사항
- Advisor: Downey, Douglas C.
- 학위논문주기
- Thesis (Ph.D.)--Northwestern University, 2024.
- 초록/해제
- 요약In this dissertation, we seek to evaluate LLM capabilities for reviewing and revising scientific documents and to develop new methods to improve them. The capabilities of large language models (LLMs) have advanced dramatically in recent years, performing on par with humans in some tasks. However, the ability of models to comprehend and produce long, highly technical text-such as that of scientific papers-remains under-explored.We construct ARIES, a dataset of scientific paper drafts, their associated peer reviews, and the new drafts after reviews, and we link individual feedback comments to specific edits that address them. Using ARIES, we study the ability of LLMs to edit scientific papers in response to feedback and to generate feedback comments.Our findings suggest that LLMs do show potential for generating feedback comments and edits for papers, but still suffer from significant limitations when attempting to comprehend or produce nuanced and technical text, often exhibiting surface-level reasoning and producing generic outputs. When revising a document in response to feedback, LLMs often write edits by quoting or paraphrasing the given feedback (48% of the time, compared to 4% for humans) and tend to include less technical detail (38% of model edits vs 53% of human edits had technical details). Similarly, when generating feedback comments for papers, baseline methods using GPT-4 were rated by users as producing generic or very generic comments more than half the time, and only 1.5 comments per paper were rated as good overall in the best baseline. We explore ways to mitigate these shortcomings and develop MARG-S, an approach for generating paper feedback using multiple specialized LLM instances that engage in internal discussion. We show that MARG-S substantially improves the ability of GPT-4 to generate specific and helpful feedback, reducing the rate of generic comments from 51% to 17% and generating 4.2 good comments per paper (a 2.8x improvement).
- 일반주제명
- Computer science
- 키워드
- Machine learning
- 키워드
- Peer review
- 기타저자
- Northwestern University Computer Science
- 기본자료저록
- Dissertations Abstracts International. 85-10B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017160152
■00520250211150920
■006m o d
■007cr#unu||||||||
■020 ▼a9798381975161
■035 ▼a(MiAaPQ)AAI30819619
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aD'Arcy, Mike.▼0(orcid)0000-0003-0355-7157
■24510▼aLarge Language Models for Automatic Peer Review and Revision in Scientific Documents
■260 ▼a[Sl]▼bNorthwestern University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a170 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 85-10, Section: B.
■500 ▼aAdvisor: Downey, Douglas C.
■5021 ▼aThesis (Ph.D.)--Northwestern University, 2024.
■520 ▼aIn this dissertation, we seek to evaluate LLM capabilities for reviewing and revising scientific documents and to develop new methods to improve them. The capabilities of large language models (LLMs) have advanced dramatically in recent years, performing on par with humans in some tasks. However, the ability of models to comprehend and produce long, highly technical text-such as that of scientific papers-remains under-explored.We construct ARIES, a dataset of scientific paper drafts, their associated peer reviews, and the new drafts after reviews, and we link individual feedback comments to specific edits that address them. Using ARIES, we study the ability of LLMs to edit scientific papers in response to feedback and to generate feedback comments.Our findings suggest that LLMs do show potential for generating feedback comments and edits for papers, but still suffer from significant limitations when attempting to comprehend or produce nuanced and technical text, often exhibiting surface-level reasoning and producing generic outputs. When revising a document in response to feedback, LLMs often write edits by quoting or paraphrasing the given feedback (48% of the time, compared to 4% for humans) and tend to include less technical detail (38% of model edits vs 53% of human edits had technical details). Similarly, when generating feedback comments for papers, baseline methods using GPT-4 were rated by users as producing generic or very generic comments more than half the time, and only 1.5 comments per paper were rated as good overall in the best baseline. We explore ways to mitigate these shortcomings and develop MARG-S, an approach for generating paper feedback using multiple specialized LLM instances that engage in internal discussion. We show that MARG-S substantially improves the ability of GPT-4 to generate specific and helpful feedback, reducing the rate of generic comments from 51% to 17% and generating 4.2 good comments per paper (a 2.8x improvement).
■590 ▼aSchool code: 0163.
■650 4▼aComputer science
■653 ▼aLanguage modeling
■653 ▼aMachine learning
■653 ▼aNatural language processing
■653 ▼aPeer review
■653 ▼aWriting assistance
■690 ▼a0800
■690 ▼a0984
■71020▼aNorthwestern University▼bComputer Science.
■7730 ▼tDissertations Abstracts International▼g85-10B.
■790 ▼a0163
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17160152▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


