본문

서브메뉴

Semi-Supervised and Representation Learning for Improved Classification and Stratification in EHR Data
Semi-Supervised and Representation Learning for Improved Classification and Stratification...
Semi-Supervised and Representation Learning for Improved Classification and Stratification in EHR Data

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202103532
ISBN  
9798280718678
DDC  
574
저자명  
Wang, Linshanshan.
서명/저자  
Semi-Supervised and Representation Learning for Improved Classification and Stratification in EHR Data
발행사항  
[Sl] : Harvard University, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
193 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-12, Section: B.
주기사항  
Advisor: Cai, Tianxi.
학위논문주기  
Thesis (Ph.D.)--Harvard University, 2025.
초록/해제  
요약The rapid digitization of healthcare has given rise to vast repositories of electronic health record (EHR) data, offering unprecedented opportunities for data-driven advancements in disease prediction, patient stratification, and clinical decision-making. However, the high dimensionality, sparsity, and heterogeneity of EHR data present unique statistical and computational challenges. Moreover, the scarcity of high-quality labels-due to the cost and complexity of manual annotation-further complicates supervised modeling efforts. This dissertation addresses these challenges through a unified framework of semi-supervised learning and representation learning for improved classification and stratification in EHR data, with applications to phenotyping, disability prediction, and patient subgroup discovery.The overarching goal of this work is to develop scalable, robust, and interpretable methods that leverage both labeled and unlabeled EHR data, improve generalizability across populations, and uncover clinically meaningful structure in complex disease settings. The dissertation is composed of three interrelated papers, each tackling a key methodological bottleneck in modern EHR-based machine learning: (1) evaluating model performance under distributional shift, (2) learning rich patient representations in the presence of limited labels, and (3) stratifying heterogeneous patient populations using outcome-informed embeddings.In Chapter 1, we consider the problem of evaluating the performance of binary classifiers when labeled data are unavailable in a target population. This setting is common in clinical phenotyping tasks, where models are trained using limited chart-reviewed labels in one cohort and then applied to other cohorts with potentially different covariate distributions. We propose STEAM (Semi-supervised Transfer lEarning of Accuracy Measures), a doubly robust estimation procedure for receiver operating characteristic (ROC) parameters under covariate shift. STEAM combines calibrated density ratio weighting with robust outcome imputation, using both unlabeled source and target data to improve efficiency while protecting against model misspecification. Through theoretical guarantees and empirical results, we demonstrate that STEAM enables accurate performance assessment in unlabeled target populations, with applications to phenotyping models in rheumatoid arthritis on temporally evolving EHR cohort.Building on the challenge of label scarcity, Chapter 2 shifts focus to semi-supervised representation learning for predictive modeling. We propose SCORE (Semi-supervised Clustering thrOugh REp-resentation learning), a generative embedding framework that models the joint distribution of high-dimensional EHR features using a multivariate Poisson-LogNormal distribution, with pretrained code embeddings capturing semantic relationships between clinical concepts. SCORE integrates limited labeled data via a hybrid Expectation-Maximization and Gaussian Variational Approximation algorithm, enabling efficient and theoretically sound inference in large-scale, partially labeled cohorts. We show that SCORE produces informative and transferable patient embeddings, improving prediction of disability status in multiple sclerosis (MS) and outperforming conventional supervised and unsupervised methods.Finally, Chapter 3 addresses the critical task of patient stratification in heterogeneous diseases. We focus on Alzheimer's disease (AD), where progression and prognosis vary substantially with age. We propose SOLAR (age-Specific Outcome-guided representation Learning for pAtient clusteRing), a novel clustering framework that incorporates time-to-event outcomes and explicitly models age-group structure using a multitask learning paradigm. SOLAR jointly learns low-dimensional patient representations across age groups, encouraging shared structure while allowing age-specific flexibility. By integrating survival information and modeling age-related heterogeneity, SOLAR identifies clinically meaningful AD subtypes with distinct prognostic profiles, improving both interpretability and clinical utility over existing age-unaware or outcome-agnostic methods.Together, these three works present a cohesive framework for semi-supervised and representation learning in EHR analysis. The methods developed here contribute new strategies for evaluating, predicting, and stratifying patient outcomes in data-scarce, high-dimensional clinical settings. In doing so, they aim to advance the broader goals of personalized medicine and evidence-based healthcare by making machine learning more robust, scalable, and clinically relevant.
일반주제명  
Biostatistics
일반주제명  
Bioinformatics
키워드  
Classification
키워드  
Representation learning
키워드  
Semi-supervised learning
키워드  
Electronic health record
기타저자  
Harvard University Biostatistics
기본자료저록  
Dissertations Abstracts International. 86-12B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017357578
■00520260202103532
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798280718678
■035    ▼a(MiAaPQ)AAI32040038
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a574
■1001  ▼aWang,  Linshanshan.▼0(orcid)0000-0002-0513-8629
■24510▼aSemi-Supervised  and  Representation  Learning  for  Improved  Classification  and  Stratification  in  EHR  Data
■260    ▼a[Sl]▼bHarvard  University▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a193  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-12,  Section:  B.
■500    ▼aAdvisor:  Cai,  Tianxi.
■5021  ▼aThesis  (Ph.D.)--Harvard  University,  2025.
■520    ▼aThe  rapid  digitization  of  healthcare  has  given  rise  to  vast  repositories  of  electronic  health  record  (EHR)  data,  offering  unprecedented  opportunities  for  data-driven  advancements  in  disease  prediction,  patient  stratification,  and  clinical  decision-making.  However,  the  high  dimensionality,  sparsity,  and  heterogeneity  of  EHR  data  present  unique  statistical  and  computational  challenges.  Moreover,  the  scarcity  of  high-quality  labels-due  to  the  cost  and  complexity  of  manual  annotation-further  complicates  supervised  modeling  efforts.  This  dissertation  addresses  these  challenges  through  a  unified  framework  of  semi-supervised  learning  and  representation  learning  for  improved  classification  and  stratification  in  EHR  data,  with  applications  to  phenotyping,  disability  prediction,  and  patient  subgroup  discovery.The  overarching  goal  of  this  work  is  to  develop  scalable,  robust,  and  interpretable  methods  that  leverage  both  labeled  and  unlabeled  EHR  data,  improve  generalizability  across  populations,  and  uncover  clinically  meaningful  structure  in  complex  disease  settings.  The  dissertation  is  composed  of  three  interrelated  papers,  each  tackling  a  key  methodological  bottleneck  in  modern  EHR-based  machine  learning:  (1)  evaluating  model  performance  under  distributional  shift,  (2)  learning  rich  patient  representations  in  the  presence  of  limited  labels,  and  (3)  stratifying  heterogeneous  patient  populations  using  outcome-informed  embeddings.In  Chapter  1,  we  consider  the  problem  of  evaluating  the  performance  of  binary  classifiers  when  labeled  data  are  unavailable  in  a  target  population.  This  setting  is  common  in  clinical  phenotyping  tasks,  where  models  are  trained  using  limited  chart-reviewed  labels  in  one  cohort  and  then  applied  to  other  cohorts  with  potentially  different  covariate  distributions.  We  propose  STEAM  (Semi-supervised  Transfer  lEarning  of  Accuracy  Measures),  a  doubly  robust  estimation  procedure  for  receiver  operating  characteristic  (ROC)  parameters  under  covariate  shift.  STEAM  combines  calibrated  density  ratio  weighting  with  robust  outcome  imputation,  using  both  unlabeled  source  and  target  data  to  improve  efficiency  while  protecting  against  model  misspecification.  Through  theoretical  guarantees  and  empirical  results,  we  demonstrate  that  STEAM  enables  accurate  performance  assessment  in  unlabeled  target  populations,  with  applications  to  phenotyping  models  in  rheumatoid  arthritis  on  temporally  evolving  EHR  cohort.Building  on  the  challenge  of  label  scarcity,  Chapter  2  shifts  focus  to  semi-supervised  representation  learning  for  predictive  modeling.  We  propose  SCORE  (Semi-supervised  Clustering  thrOugh  REp-resentation  learning),  a  generative  embedding  framework  that  models  the  joint  distribution  of  high-dimensional  EHR  features  using  a  multivariate  Poisson-LogNormal  distribution,  with  pretrained  code  embeddings  capturing  semantic  relationships  between  clinical  concepts.  SCORE  integrates  limited  labeled  data  via  a  hybrid  Expectation-Maximization  and  Gaussian  Variational  Approximation  algorithm,  enabling  efficient  and  theoretically  sound  inference  in  large-scale,  partially  labeled  cohorts.  We  show  that  SCORE  produces  informative  and  transferable  patient  embeddings,  improving  prediction  of  disability  status  in  multiple  sclerosis  (MS)  and  outperforming  conventional  supervised  and  unsupervised  methods.Finally,  Chapter  3  addresses  the  critical  task  of  patient  stratification  in  heterogeneous  diseases.  We  focus  on  Alzheimer's  disease  (AD),  where  progression  and  prognosis  vary  substantially  with  age.  We  propose  SOLAR  (age-Specific  Outcome-guided  representation  Learning  for  pAtient  clusteRing),  a  novel  clustering  framework  that  incorporates  time-to-event  outcomes  and  explicitly  models  age-group  structure  using  a  multitask  learning  paradigm.  SOLAR  jointly  learns  low-dimensional  patient  representations  across  age  groups,  encouraging  shared  structure  while  allowing  age-specific  flexibility.  By  integrating  survival  information  and  modeling  age-related  heterogeneity,  SOLAR  identifies  clinically  meaningful  AD  subtypes  with  distinct  prognostic  profiles,  improving  both  interpretability  and  clinical  utility  over  existing  age-unaware  or  outcome-agnostic  methods.Together,  these  three  works  present  a  cohesive  framework  for  semi-supervised  and  representation  learning  in  EHR  analysis.  The  methods  developed  here  contribute  new  strategies  for  evaluating,  predicting,  and  stratifying  patient  outcomes  in  data-scarce,  high-dimensional  clinical  settings.  In  doing  so,  they  aim  to  advance  the  broader  goals  of  personalized  medicine  and  evidence-based  healthcare  by  making  machine  learning  more  robust,  scalable,  and  clinically  relevant.
■590    ▼aSchool  code:  0084.
■650  4▼aBiostatistics
■650  4▼aBioinformatics
■653    ▼aClassification
■653    ▼aRepresentation  learning
■653    ▼aSemi-supervised  learning
■653    ▼aElectronic  health  record
■690    ▼a0308
■690    ▼a0769
■690    ▼a0715
■71020▼aHarvard  University▼bBiostatistics.
■7730  ▼tDissertations  Abstracts  International▼g86-12B.
■790    ▼a0084
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357578▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF17859 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.