서브메뉴
검색
Semi-Supervised and Representation Learning for Improved Classification and Stratification in EHR Data
Semi-Supervised and Representation Learning for Improved Classification and Stratification in EHR Data
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103532
- ISBN
- 9798280718678
- DDC
- 574
- 서명/저자
- Semi-Supervised and Representation Learning for Improved Classification and Stratification in EHR Data
- 발행사항
- [Sl] : Harvard University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 193 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-12, Section: B.
- 주기사항
- Advisor: Cai, Tianxi.
- 학위논문주기
- Thesis (Ph.D.)--Harvard University, 2025.
- 초록/해제
- 요약The rapid digitization of healthcare has given rise to vast repositories of electronic health record (EHR) data, offering unprecedented opportunities for data-driven advancements in disease prediction, patient stratification, and clinical decision-making. However, the high dimensionality, sparsity, and heterogeneity of EHR data present unique statistical and computational challenges. Moreover, the scarcity of high-quality labels-due to the cost and complexity of manual annotation-further complicates supervised modeling efforts. This dissertation addresses these challenges through a unified framework of semi-supervised learning and representation learning for improved classification and stratification in EHR data, with applications to phenotyping, disability prediction, and patient subgroup discovery.The overarching goal of this work is to develop scalable, robust, and interpretable methods that leverage both labeled and unlabeled EHR data, improve generalizability across populations, and uncover clinically meaningful structure in complex disease settings. The dissertation is composed of three interrelated papers, each tackling a key methodological bottleneck in modern EHR-based machine learning: (1) evaluating model performance under distributional shift, (2) learning rich patient representations in the presence of limited labels, and (3) stratifying heterogeneous patient populations using outcome-informed embeddings.In Chapter 1, we consider the problem of evaluating the performance of binary classifiers when labeled data are unavailable in a target population. This setting is common in clinical phenotyping tasks, where models are trained using limited chart-reviewed labels in one cohort and then applied to other cohorts with potentially different covariate distributions. We propose STEAM (Semi-supervised Transfer lEarning of Accuracy Measures), a doubly robust estimation procedure for receiver operating characteristic (ROC) parameters under covariate shift. STEAM combines calibrated density ratio weighting with robust outcome imputation, using both unlabeled source and target data to improve efficiency while protecting against model misspecification. Through theoretical guarantees and empirical results, we demonstrate that STEAM enables accurate performance assessment in unlabeled target populations, with applications to phenotyping models in rheumatoid arthritis on temporally evolving EHR cohort.Building on the challenge of label scarcity, Chapter 2 shifts focus to semi-supervised representation learning for predictive modeling. We propose SCORE (Semi-supervised Clustering thrOugh REp-resentation learning), a generative embedding framework that models the joint distribution of high-dimensional EHR features using a multivariate Poisson-LogNormal distribution, with pretrained code embeddings capturing semantic relationships between clinical concepts. SCORE integrates limited labeled data via a hybrid Expectation-Maximization and Gaussian Variational Approximation algorithm, enabling efficient and theoretically sound inference in large-scale, partially labeled cohorts. We show that SCORE produces informative and transferable patient embeddings, improving prediction of disability status in multiple sclerosis (MS) and outperforming conventional supervised and unsupervised methods.Finally, Chapter 3 addresses the critical task of patient stratification in heterogeneous diseases. We focus on Alzheimer's disease (AD), where progression and prognosis vary substantially with age. We propose SOLAR (age-Specific Outcome-guided representation Learning for pAtient clusteRing), a novel clustering framework that incorporates time-to-event outcomes and explicitly models age-group structure using a multitask learning paradigm. SOLAR jointly learns low-dimensional patient representations across age groups, encouraging shared structure while allowing age-specific flexibility. By integrating survival information and modeling age-related heterogeneity, SOLAR identifies clinically meaningful AD subtypes with distinct prognostic profiles, improving both interpretability and clinical utility over existing age-unaware or outcome-agnostic methods.Together, these three works present a cohesive framework for semi-supervised and representation learning in EHR analysis. The methods developed here contribute new strategies for evaluating, predicting, and stratifying patient outcomes in data-scarce, high-dimensional clinical settings. In doing so, they aim to advance the broader goals of personalized medicine and evidence-based healthcare by making machine learning more robust, scalable, and clinically relevant.
- 일반주제명
- Biostatistics
- 일반주제명
- Bioinformatics
- 키워드
- Classification
- 기타저자
- Harvard University Biostatistics
- 기본자료저록
- Dissertations Abstracts International. 86-12B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017357578
■00520260202103532
■006m o d
■007cr#unu||||||||
■020 ▼a9798280718678
■035 ▼a(MiAaPQ)AAI32040038
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a574
■1001 ▼aWang, Linshanshan.▼0(orcid)0000-0002-0513-8629
■24510▼aSemi-Supervised and Representation Learning for Improved Classification and Stratification in EHR Data
■260 ▼a[Sl]▼bHarvard University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a193 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-12, Section: B.
■500 ▼aAdvisor: Cai, Tianxi.
■5021 ▼aThesis (Ph.D.)--Harvard University, 2025.
■520 ▼aThe rapid digitization of healthcare has given rise to vast repositories of electronic health record (EHR) data, offering unprecedented opportunities for data-driven advancements in disease prediction, patient stratification, and clinical decision-making. However, the high dimensionality, sparsity, and heterogeneity of EHR data present unique statistical and computational challenges. Moreover, the scarcity of high-quality labels-due to the cost and complexity of manual annotation-further complicates supervised modeling efforts. This dissertation addresses these challenges through a unified framework of semi-supervised learning and representation learning for improved classification and stratification in EHR data, with applications to phenotyping, disability prediction, and patient subgroup discovery.The overarching goal of this work is to develop scalable, robust, and interpretable methods that leverage both labeled and unlabeled EHR data, improve generalizability across populations, and uncover clinically meaningful structure in complex disease settings. The dissertation is composed of three interrelated papers, each tackling a key methodological bottleneck in modern EHR-based machine learning: (1) evaluating model performance under distributional shift, (2) learning rich patient representations in the presence of limited labels, and (3) stratifying heterogeneous patient populations using outcome-informed embeddings.In Chapter 1, we consider the problem of evaluating the performance of binary classifiers when labeled data are unavailable in a target population. This setting is common in clinical phenotyping tasks, where models are trained using limited chart-reviewed labels in one cohort and then applied to other cohorts with potentially different covariate distributions. We propose STEAM (Semi-supervised Transfer lEarning of Accuracy Measures), a doubly robust estimation procedure for receiver operating characteristic (ROC) parameters under covariate shift. STEAM combines calibrated density ratio weighting with robust outcome imputation, using both unlabeled source and target data to improve efficiency while protecting against model misspecification. Through theoretical guarantees and empirical results, we demonstrate that STEAM enables accurate performance assessment in unlabeled target populations, with applications to phenotyping models in rheumatoid arthritis on temporally evolving EHR cohort.Building on the challenge of label scarcity, Chapter 2 shifts focus to semi-supervised representation learning for predictive modeling. We propose SCORE (Semi-supervised Clustering thrOugh REp-resentation learning), a generative embedding framework that models the joint distribution of high-dimensional EHR features using a multivariate Poisson-LogNormal distribution, with pretrained code embeddings capturing semantic relationships between clinical concepts. SCORE integrates limited labeled data via a hybrid Expectation-Maximization and Gaussian Variational Approximation algorithm, enabling efficient and theoretically sound inference in large-scale, partially labeled cohorts. We show that SCORE produces informative and transferable patient embeddings, improving prediction of disability status in multiple sclerosis (MS) and outperforming conventional supervised and unsupervised methods.Finally, Chapter 3 addresses the critical task of patient stratification in heterogeneous diseases. We focus on Alzheimer's disease (AD), where progression and prognosis vary substantially with age. We propose SOLAR (age-Specific Outcome-guided representation Learning for pAtient clusteRing), a novel clustering framework that incorporates time-to-event outcomes and explicitly models age-group structure using a multitask learning paradigm. SOLAR jointly learns low-dimensional patient representations across age groups, encouraging shared structure while allowing age-specific flexibility. By integrating survival information and modeling age-related heterogeneity, SOLAR identifies clinically meaningful AD subtypes with distinct prognostic profiles, improving both interpretability and clinical utility over existing age-unaware or outcome-agnostic methods.Together, these three works present a cohesive framework for semi-supervised and representation learning in EHR analysis. The methods developed here contribute new strategies for evaluating, predicting, and stratifying patient outcomes in data-scarce, high-dimensional clinical settings. In doing so, they aim to advance the broader goals of personalized medicine and evidence-based healthcare by making machine learning more robust, scalable, and clinically relevant.
■590 ▼aSchool code: 0084.
■650 4▼aBiostatistics
■650 4▼aBioinformatics
■653 ▼aClassification
■653 ▼aRepresentation learning
■653 ▼aSemi-supervised learning
■653 ▼aElectronic health record
■690 ▼a0308
■690 ▼a0769
■690 ▼a0715
■71020▼aHarvard University▼bBiostatistics.
■7730 ▼tDissertations Abstracts International▼g86-12B.
■790 ▼a0084
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357578▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


