서브메뉴
검색
Detecting Risky Alcohol Use With Natural Language Processing and Computable Phenotypes in Clinical Records
Detecting Risky Alcohol Use With Natural Language Processing and Computable Phenotypes in Clinical Records
Detailed Information
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211153007
- ISBN
- 9798384044314
- DDC
- 004
- 서명/저자
- Detecting Risky Alcohol Use With Natural Language Processing and Computable Phenotypes in Clinical Records
- 발행사항
- [Sl] : University of Michigan, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 141 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-03, Section: A.
- 주기사항
- Advisor: Vydiswaran, V. G. Vinod.
- 학위논문주기
- Thesis (Ph.D.)--University of Michigan, 2024.
- 초록/해제
- 요약Alcohol use is common throughout the United States. Alcohol Use Disorder (AUD) is estimated to affect 12-14% of the US population, and has destructive impacts on the physical and social health of an individual. Unless a person has received diagnosis or treatment for AUD, information about their consumption is largely restricted to free-text notes in their social history and is difficult to locate with simple searches. Because of the correlation between alcohol use and poor surgical outcomes, there is a need to locate alcohol-related information and calculate alcohol-use risk for clinicians who may be seeing a patient for the first time before a procedure. This information alert the preoperative team to this under-recognized surgical risk factor, triggering additional alcohol screening that could aid in clinical decision making.The aims of this dissertation are to develop a natural language processing (NLP) classifier that assesses the text in health records to issue a label representing the degree to which the patient experiences risky alcohol use; and to develop a computable phenotype for risky alcohol use that uses structured data in the clinical record.A binary (high-risk/not high-risk) NLP classifier has an F1 score of 0.78, far outperforming the ability of ICD codes alone to correctly identify high-risk patients. An ordered four-class NLP algorithm applies a transformer encoder and a CNN for intermediate labeling and a bidirectional LSTM neural network as an inference head to effectively build a model with scarce data in rare classes to an overall macro F1 score of 0.77, with true-positive performance of 0.83 and 0.73 for Probable-Dependence and High-Risk, respectively. The proposed structured-data computable phenotype is novel in its ability to provide stratified levels of risk and as a two-class classifier, is significantly more effective than others in the literature. However, it is unable to classify 41% of patients and relies on non-standard structured data in the record for its improvements over another published phenotype.Finally, we consider this model in the context of other Large Language Model approaches to clinical concept extraction, examine the utility of alternatives to the F1 statistic for model selection in ordinal classifiers, and validate the NLP approach to extracting and calculating a numeric value for a patient's weekly alcohol consumption. This dissertation's contributions include a multi-stage, novel approach for extracting sparse information in a noisy, imbalanced dataset, a new ordinal NLP classifier representing alcohol-use risk, and a four-class computable phenotype for alcohol-use risk.
- 일반주제명
- Information technology
- 일반주제명
- American studies
- 일반주제명
- Psychology
- 일반주제명
- Mental health
- 키워드
- United States
- 키워드
- US population
- 키워드
- ICD codes
- 기타저자
- University of Michigan Hlth Infrastr & Lrng Systs PhD
- 기본자료저록
- Dissertations Abstracts International. 86-03A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017164478
■00520250211153007
■006m o d
■007cr#unu||||||||
■020 ▼a9798384044314
■035 ▼a(MiAaPQ)AAI31631407
■035 ▼a(MiAaPQ)umichrackham005726
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aWeber, Katherine G.
■24510▼aDetecting Risky Alcohol Use With Natural Language Processing and Computable Phenotypes in Clinical Records
■260 ▼a[Sl]▼bUniversity of Michigan▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a141 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-03, Section: A.
■500 ▼aAdvisor: Vydiswaran, V. G. Vinod.
■5021 ▼aThesis (Ph.D.)--University of Michigan, 2024.
■520 ▼aAlcohol use is common throughout the United States. Alcohol Use Disorder (AUD) is estimated to affect 12-14% of the US population, and has destructive impacts on the physical and social health of an individual. Unless a person has received diagnosis or treatment for AUD, information about their consumption is largely restricted to free-text notes in their social history and is difficult to locate with simple searches. Because of the correlation between alcohol use and poor surgical outcomes, there is a need to locate alcohol-related information and calculate alcohol-use risk for clinicians who may be seeing a patient for the first time before a procedure. This information alert the preoperative team to this under-recognized surgical risk factor, triggering additional alcohol screening that could aid in clinical decision making.The aims of this dissertation are to develop a natural language processing (NLP) classifier that assesses the text in health records to issue a label representing the degree to which the patient experiences risky alcohol use; and to develop a computable phenotype for risky alcohol use that uses structured data in the clinical record.A binary (high-risk/not high-risk) NLP classifier has an F1 score of 0.78, far outperforming the ability of ICD codes alone to correctly identify high-risk patients. An ordered four-class NLP algorithm applies a transformer encoder and a CNN for intermediate labeling and a bidirectional LSTM neural network as an inference head to effectively build a model with scarce data in rare classes to an overall macro F1 score of 0.77, with true-positive performance of 0.83 and 0.73 for Probable-Dependence and High-Risk, respectively. The proposed structured-data computable phenotype is novel in its ability to provide stratified levels of risk and as a two-class classifier, is significantly more effective than others in the literature. However, it is unable to classify 41% of patients and relies on non-standard structured data in the record for its improvements over another published phenotype.Finally, we consider this model in the context of other Large Language Model approaches to clinical concept extraction, examine the utility of alternatives to the F1 statistic for model selection in ordinal classifiers, and validate the NLP approach to extracting and calculating a numeric value for a patient's weekly alcohol consumption. This dissertation's contributions include a multi-stage, novel approach for extracting sparse information in a noisy, imbalanced dataset, a new ordinal NLP classifier representing alcohol-use risk, and a four-class computable phenotype for alcohol-use risk.
■590 ▼aSchool code: 0127.
■650 4▼aInformation technology
■650 4▼aAmerican studies
■650 4▼aPsychology
■650 4▼aMental health
■653 ▼aNatural language processing
■653 ▼aAlcohol Use Disorder
■653 ▼aUnited States
■653 ▼aUS population
■653 ▼aICD codes
■690 ▼a0489
■690 ▼a0323
■690 ▼a0621
■690 ▼a0347
■71020▼aUniversity of Michigan▼bHlth Infrastr & Lrng Systs PhD.
■7730 ▼tDissertations Abstracts International▼g86-03A.
■790 ▼a0127
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164478▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.
Preview
Export
ChatGPT Discussion
AI Recommended Related Books
Buch Status
- Reservierung
- frei buchen
- Meine Mappe
- Erste Aufräumarbeiten Anfrage
- Non-Book Loan Application
- Nighttime Book Loan Application
Available after logging in.


