서브메뉴
검색
The Development of a Novel Large Language Model Method for Identification of Incarceration History via the Electronic Health Record and Evaluation of Care Processes in the Emergency Department Setting
The Development of a Novel Large Language Model Method for Identification of Incarceration History via the Electronic Health Record and Evaluation of Care Processes in the Emergency Department Setting
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103052
- ISBN
- 9798314887721
- DDC
- 610
- 저자명
- Huang, Thomas.
- 서명/저자
- The Development of a Novel Large Language Model Method for Identification of Incarceration History via the Electronic Health Record and Evaluation of Care Processes in the Emergency Department Setting
- 발행사항
- [Sl] : Yale University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 66 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-11, Section: B.
- 주기사항
- Advisor: Taylor, R. Andrew.
- 학위논문주기
- Thesis (M.D.)--Yale University, 2025.
- 초록/해제
- 요약Objective: Incarceration is a significant social driver of health and patients with a history of incarceration face systemic healthcare disparities, higher morbidity, mortality, and racialized health inequities. Incarceration status is largely invisible due to poor electronic health record (EHR) capture. In this thesis, we aim to develop, train, and validate a novel natural language processing technique to more effectively identify incarceration status in the EHR and apply this method in the emergency department setting to demonstrate proof of concept and elucidate care process disparities.Methods: The study population consisted of adult patients (≥ 18 y.o.) who presented to the emergency department between June 2013 and August 2021. The EHR database was filtered for notes for specific incarceration-related-terms, and then a random selection of 1,000 notes were annotated for incarceration and further stratified into specific statuses of prior history, recent, and current incarceration. For natural language processing (NLP) model development, 80% of the notes were used to train the Longformer-based and RoBERTa algorithms. The remaining 20% of the notes underwent analysis with GPT-4. The fine-tuned Clinical-Longformer model was subsequently applied to 480,374 notes from the ED setting. Socio-demographics, co-morbidities, and care processes were compared between patients with and without history of incarceration as identified by the LLM. We utilized a multivariable logistic regression to assess independent correlation of incarceration history and care processes in the ED.Results: Manual annotation revealed that 559 of 1000 notes (55.9%) contained evidence of incarceration history. ICD-10 code (sensitivity: 4.8%, specificity: 99.1%, F1-score: 0.09) demonstrated inferior performance to RoBERTa NLP (sensitivity: 78.6%, specificity: 73.3%, F1-score: 0.79), Longformer NLP (sensitivity: 94.6%, specificity: 87.5%, F1-score: 0.93) and GPT-4 (sensitivity: 100%, specificity: 61.1%, F1-score: 0.86). In a separate cohort of 177,987 ED encounters, 1,734 involved patients with a history of incarceration. These patients were more likely to be male, Black, Hispanic, or of other race/ethnicity, unemployed or disabled, and have smoking or substance use histories. Compared to those without incarceration histories, they had higher odds of eloping (OR: 3.59 [2.41-5.12]), leaving AMA (OR: 2.39 [1.46-3.67]), and being subjected to sedation (OR: 3.89 [3.19-4.70]) and restraints (OR: 3.76 [3.06- 4.57]). After adjusting for covariates, only the association with elopement remained significant (aOR: 1.65 [1.08-2.43]).Conclusions: Our advanced LLM demonstrates a high degree of accuracy in identifying incarceration status from clinical notes. Leveraging this method to identify highly representative cohorts of patients with history of incarceration presenting to the ED highlights the feasibility of NLP methods for means of identification. This method delineates differences in ED patient characteristics and care processes for individuals with incarceration histories, underscoring the utility of NLP in uncovering care disparities in underserved and stigmatized populations.
- 일반주제명
- Medicine
- 일반주제명
- Health sciences
- 키워드
- Incarceration
- 기타저자
- Yale University Yale School of Medicine
- 기본자료저록
- Dissertations Abstracts International. 86-11B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017356869
■00520260202103052
■006m o d
■007cr#unu||||||||
■020 ▼a9798314887721
■035 ▼a(MiAaPQ)AAI31931989
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a610
■1001 ▼aHuang, Thomas.
■24510▼aThe Development of a Novel Large Language Model Method for Identification of Incarceration History via the Electronic Health Record and Evaluation of Care Processes in the Emergency Department Setting
■260 ▼a[Sl]▼bYale University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a66 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-11, Section: B.
■500 ▼aAdvisor: Taylor, R. Andrew.
■5021 ▼aThesis (M.D.)--Yale University, 2025.
■520 ▼aObjective: Incarceration is a significant social driver of health and patients with a history of incarceration face systemic healthcare disparities, higher morbidity, mortality, and racialized health inequities. Incarceration status is largely invisible due to poor electronic health record (EHR) capture. In this thesis, we aim to develop, train, and validate a novel natural language processing technique to more effectively identify incarceration status in the EHR and apply this method in the emergency department setting to demonstrate proof of concept and elucidate care process disparities.Methods: The study population consisted of adult patients (≥ 18 y.o.) who presented to the emergency department between June 2013 and August 2021. The EHR database was filtered for notes for specific incarceration-related-terms, and then a random selection of 1,000 notes were annotated for incarceration and further stratified into specific statuses of prior history, recent, and current incarceration. For natural language processing (NLP) model development, 80% of the notes were used to train the Longformer-based and RoBERTa algorithms. The remaining 20% of the notes underwent analysis with GPT-4. The fine-tuned Clinical-Longformer model was subsequently applied to 480,374 notes from the ED setting. Socio-demographics, co-morbidities, and care processes were compared between patients with and without history of incarceration as identified by the LLM. We utilized a multivariable logistic regression to assess independent correlation of incarceration history and care processes in the ED.Results: Manual annotation revealed that 559 of 1000 notes (55.9%) contained evidence of incarceration history. ICD-10 code (sensitivity: 4.8%, specificity: 99.1%, F1-score: 0.09) demonstrated inferior performance to RoBERTa NLP (sensitivity: 78.6%, specificity: 73.3%, F1-score: 0.79), Longformer NLP (sensitivity: 94.6%, specificity: 87.5%, F1-score: 0.93) and GPT-4 (sensitivity: 100%, specificity: 61.1%, F1-score: 0.86). In a separate cohort of 177,987 ED encounters, 1,734 involved patients with a history of incarceration. These patients were more likely to be male, Black, Hispanic, or of other race/ethnicity, unemployed or disabled, and have smoking or substance use histories. Compared to those without incarceration histories, they had higher odds of eloping (OR: 3.59 [2.41-5.12]), leaving AMA (OR: 2.39 [1.46-3.67]), and being subjected to sedation (OR: 3.89 [3.19-4.70]) and restraints (OR: 3.76 [3.06- 4.57]). After adjusting for covariates, only the association with elopement remained significant (aOR: 1.65 [1.08-2.43]).Conclusions: Our advanced LLM demonstrates a high degree of accuracy in identifying incarceration status from clinical notes. Leveraging this method to identify highly representative cohorts of patients with history of incarceration presenting to the ED highlights the feasibility of NLP methods for means of identification. This method delineates differences in ED patient characteristics and care processes for individuals with incarceration histories, underscoring the utility of NLP in uncovering care disparities in underserved and stigmatized populations.
■590 ▼aSchool code: 0265.
■650 4▼aMedicine
■650 4▼aHealth sciences
■653 ▼aEmergency medicine
■653 ▼aHealth disparities
■653 ▼aIncarceration
■653 ▼aLarge language model
■653 ▼aNatural language processing
■690 ▼a0564
■690 ▼a0800
■690 ▼a0566
■71020▼aYale University▼bYale School of Medicine.
■7730 ▼tDissertations Abstracts International▼g86-11B.
■790 ▼a0265
■791 ▼aM.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17356869▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


