본문

서브메뉴

Causal Inference in Complex Observational Settings With Applications to Electronic Health Record Data
Causal Inference in Complex Observational Settings With Applications to Electronic Health ...
Causal Inference in Complex Observational Settings With Applications to Electronic Health Record Data

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202103543
ISBN  
9798280716261
DDC  
574
저자명  
Xu, Daniel.
서명/저자  
Causal Inference in Complex Observational Settings With Applications to Electronic Health Record Data
발행사항  
[Sl] : Harvard University, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
208 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-12, Section: B.
주기사항  
Advisor: Mukherjee, Rajarshi.
학위논문주기  
Thesis (Ph.D.)--Harvard University, 2025.
초록/해제  
요약Electronic health record (EHR) data serve as a valuable source of real-world evidence for assessing treatment effects, offering rich, longitudinal patient data that can enhance clinical research, inform healthcare decisions, and ultimately improve patient outcomes. However, leveraging these complex data sources for reliable statistical inference remains challenging. For example, one common difficulty is the limited availability of readily validated clinical outcomes in EHR data, which often requires labor-intensive manual annotation through chart review. A related challenge arises when studies span multiple institutions, where differences in patient populations can introduce heterogeneity and potential biases. These complications further exacerbate the fundamental challenge of using observational data to draw valid causal conclusions. In this dissertation, we propose novel statistical methods for causal inference that address certain complexities relevant to clinical studies leveraging EHR data. In particular, we focus on two settings of interest: (1) semi-supervised learning, where outcome labels are scarce, and (2) transfer learning, where data must be integrated across diverse populations. We hope that the contributions of this dissertation offer meaningful steps toward unlocking the full potential of EHR data for clinical and translational research.In Chapter 1, we address the problem of estimating treatment effects in a semi-supervised setting where labeled outcomes are sparse but unlabeled data are abundant. We develop a semi-supervised calibration method that leverages a subsample of labeled outcomes to calibrate inferred outcomes, ensuring that the downstream treatment effect estimator remains consistent despite potential errors in outcome imputation. Unlike most traditional semi-supervised methods, we allow the labeling mechanism to depend on the observed data, rather than assuming it is completely random. This problem is analogous to estimating mean outcomes in longitudinal studies with monotone missingness, and we show that our proposed estimator is asymptotically equivalent to the augmented inverse probability weighting (AIPW) estimator when a consistent estimate of the labeling propensity score is available. The estimator is multiply robust and locally semiparametric efficient. We also demonstrate improved finite-sample efficiency in semi-supervised settings, owing to an effective normalization of an implicit augmentation term. The finite-sample performance is evaluated through simulations, and we illustrate the method in a case study comparing the effectiveness of two anti-TNF therapies on remission outcomes in patients with rheumatoid arthritis.In Chapter 2, we consider the problem of causal mediation analysis in a semi-supervised setting. Causal mediation analysis is a fundamental tool for understanding how treatments or exposures affect outcomes through intermediate variables. However, existing methods lack theoretical and practical guarantees in the presence of substantial missingness in the outcome variable. To address this, we propose a robust and efficient method for the semi-supervised estimation of natural direct and indirect effects, accommodating settings both with and without surrogate outcomes. Our approach extends the double machine learning framework by constructing multiply robust one-step estimators, which enable the use of flexible machine learning algorithms for estimating nuisance functions. We show that the proposed estimators are asymptotically normal and locally minimax optimal under semi-supervised models that do not assume the distribution of the non-missing variables to be known. Through extensive simulations, we demonstrate that the estimators remain unbiased, yield valid inference, and achieve substantial efficiency gains by leveraging predictive surrogates, even under complex data-generating mechanisms. We illustrate the method using EHR data to examine whether racial disparities in Alzheimer's disease-related outcomes are mediated through cardiovascular comorbidities.In Chapter 3, we explore the problem of estimating low-dimensional functionals in a target population using data from both source and target domains, where the two distributions are only weakly aligned. This setting arises frequently in practice when performing transfer learning, yet remains underexplored in the context of functional estimation. We focus on three canonical functionals relevant to statistics and causal inference: the quadratic regression functional, the expected conditional covariance, and the mean response in missing data models. Rather than making strong parametric assumptions on the source and target distributions, we model weak alignment by assuming that the differences in conditional mean functions across domains are bounded in L2 norm. For each functional, we propose two estimation strategies: (1) an influence function-based estimator that employs existing minimax-optimal transfer learning methods for nuisance estimation, and (2) a functional-level confidence thresholding estimator that selects between source- and target-based estimators. We derive non-asymptotic upper bounds on the estimation risk in mean absolute error and show that both strategies achieve faster rates than target-only estimators, provided that posterior drift between domains is sufficiently small. When functionals involve two nuisances, we demonstrate that some of the proposed estimators are able to outperform the minimax-optimal target-only estimator as long as the degree of posterior drift in either of the two nuisance functions is sufficiently small, which we refer to as distributional shift double robustness. Furthermore, we demonstrate that under additional assumptions -- for example, by directly bounding the separation between the functional evaluated at source and target distributions -- the functional-level confidence thresholding estimator can attain minimax-optimal rates up to log factors.
일반주제명  
Biostatistics
일반주제명  
Bioinformatics
일반주제명  
Information technology
키워드  
Causal inference
키워드  
Electronic health records
키워드  
Semi-supervised learning
키워드  
Transfer learning
기타저자  
Harvard University Biostatistics
기본자료저록  
Dissertations Abstracts International. 86-12B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017357664
■00520260202103543
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798280716261
■035    ▼a(MiAaPQ)AAI32041071
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a574
■1001  ▼aXu,  Daniel.▼0(orcid)0000-0002-4750-7138
■24510▼aCausal  Inference  in  Complex  Observational  Settings  With  Applications  to  Electronic  Health  Record  Data
■260    ▼a[Sl]▼bHarvard  University▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a208  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-12,  Section:  B.
■500    ▼aAdvisor:  Mukherjee,  Rajarshi.
■5021  ▼aThesis  (Ph.D.)--Harvard  University,  2025.
■520    ▼aElectronic  health  record  (EHR)  data  serve  as  a  valuable  source  of  real-world  evidence  for  assessing  treatment  effects,  offering  rich,  longitudinal  patient  data  that  can  enhance  clinical  research,  inform  healthcare  decisions,  and  ultimately  improve  patient  outcomes.  However,  leveraging  these  complex  data  sources  for  reliable  statistical  inference  remains  challenging.  For  example,  one  common  difficulty  is  the  limited  availability  of  readily  validated  clinical  outcomes  in  EHR  data,  which  often  requires  labor-intensive  manual  annotation  through  chart  review.  A  related  challenge  arises  when  studies  span  multiple  institutions,  where  differences  in  patient  populations  can  introduce  heterogeneity  and  potential  biases.  These  complications  further  exacerbate  the  fundamental  challenge  of  using  observational  data  to  draw  valid  causal  conclusions.  In  this  dissertation,  we  propose  novel  statistical  methods  for  causal  inference  that  address  certain  complexities  relevant  to  clinical  studies  leveraging  EHR  data.  In  particular,  we  focus  on  two  settings  of  interest:  (1)  semi-supervised  learning,  where  outcome  labels  are  scarce,  and  (2)  transfer  learning,  where  data  must  be  integrated  across  diverse  populations.  We  hope  that  the  contributions  of  this  dissertation  offer  meaningful  steps  toward  unlocking  the  full  potential  of  EHR  data  for  clinical  and  translational  research.In  Chapter  1,  we  address  the  problem  of  estimating  treatment  effects  in  a  semi-supervised  setting  where  labeled  outcomes  are  sparse  but  unlabeled  data  are  abundant.  We  develop  a  semi-supervised  calibration  method  that  leverages  a  subsample  of  labeled  outcomes  to  calibrate  inferred  outcomes,  ensuring  that  the  downstream  treatment  effect  estimator  remains  consistent  despite  potential  errors  in  outcome  imputation.  Unlike  most  traditional  semi-supervised  methods,  we  allow  the  labeling  mechanism  to  depend  on  the  observed  data,  rather  than  assuming  it  is  completely  random.  This  problem  is  analogous  to  estimating  mean  outcomes  in  longitudinal  studies  with  monotone  missingness,  and  we  show  that  our  proposed  estimator  is  asymptotically  equivalent  to  the  augmented  inverse  probability  weighting  (AIPW)  estimator  when  a  consistent  estimate  of  the  labeling  propensity  score  is  available.  The  estimator  is  multiply  robust  and  locally  semiparametric  efficient.  We  also  demonstrate  improved  finite-sample  efficiency  in  semi-supervised  settings,  owing  to  an  effective  normalization  of  an  implicit  augmentation  term.  The  finite-sample  performance  is  evaluated  through  simulations,  and  we  illustrate  the  method  in  a  case  study  comparing  the  effectiveness  of  two  anti-TNF  therapies  on  remission  outcomes  in  patients  with  rheumatoid  arthritis.In  Chapter  2,  we  consider  the  problem  of  causal  mediation  analysis  in  a  semi-supervised  setting.  Causal  mediation  analysis  is  a  fundamental  tool  for  understanding  how  treatments  or  exposures  affect  outcomes  through  intermediate  variables.  However,  existing  methods  lack  theoretical  and  practical  guarantees  in  the  presence  of  substantial  missingness  in  the  outcome  variable.  To  address  this,  we  propose  a  robust  and  efficient  method  for  the  semi-supervised  estimation  of  natural  direct  and  indirect  effects,  accommodating  settings  both  with  and  without  surrogate  outcomes.  Our  approach  extends  the  double  machine  learning  framework  by  constructing  multiply  robust  one-step  estimators,  which  enable  the  use  of  flexible  machine  learning  algorithms  for  estimating  nuisance  functions.  We  show  that  the  proposed  estimators  are  asymptotically  normal  and  locally  minimax  optimal  under  semi-supervised  models  that  do  not  assume  the  distribution  of  the  non-missing  variables  to  be  known.  Through  extensive  simulations,  we  demonstrate  that  the  estimators  remain  unbiased,  yield  valid  inference,  and  achieve  substantial  efficiency  gains  by  leveraging  predictive  surrogates,  even  under  complex  data-generating  mechanisms.  We  illustrate  the  method  using  EHR  data  to  examine  whether  racial  disparities  in  Alzheimer's  disease-related  outcomes  are  mediated  through  cardiovascular  comorbidities.In  Chapter  3,  we  explore  the  problem  of  estimating  low-dimensional  functionals  in  a  target  population  using  data  from  both  source  and  target  domains,  where  the  two  distributions  are  only  weakly  aligned.  This  setting  arises  frequently  in  practice  when  performing  transfer  learning,  yet  remains  underexplored  in  the  context  of  functional  estimation.  We  focus  on  three  canonical  functionals  relevant  to  statistics  and  causal  inference:  the  quadratic  regression  functional,  the  expected  conditional  covariance,  and  the  mean  response  in  missing  data  models.  Rather  than  making  strong  parametric  assumptions  on  the  source  and  target  distributions,  we  model  weak  alignment  by  assuming  that  the  differences  in  conditional  mean  functions  across  domains  are  bounded  in  L2  norm.  For  each  functional,  we  propose  two  estimation  strategies:  (1)  an  influence  function-based  estimator  that  employs  existing  minimax-optimal  transfer  learning  methods  for  nuisance  estimation,  and  (2)  a  functional-level  confidence  thresholding  estimator  that  selects  between  source-  and  target-based  estimators.  We  derive  non-asymptotic  upper  bounds  on  the  estimation  risk  in  mean  absolute  error  and  show  that  both  strategies  achieve  faster  rates  than  target-only  estimators,  provided  that  posterior  drift  between  domains  is  sufficiently  small.  When  functionals  involve  two  nuisances,  we  demonstrate  that  some  of  the  proposed  estimators  are  able  to  outperform  the  minimax-optimal  target-only  estimator  as  long  as  the  degree  of  posterior  drift  in  either  of  the  two  nuisance  functions  is  sufficiently  small,  which  we  refer  to  as  distributional  shift  double  robustness.  Furthermore,  we  demonstrate  that  under  additional  assumptions  --  for  example,  by  directly  bounding  the  separation  between  the  functional  evaluated  at  source  and  target  distributions  --  the  functional-level  confidence  thresholding  estimator  can  attain  minimax-optimal  rates  up  to  log  factors.
■590    ▼aSchool  code:  0084.
■650  4▼aBiostatistics
■650  4▼aBioinformatics
■650  4▼aInformation  technology
■653    ▼aCausal  inference
■653    ▼aElectronic  health  records
■653    ▼aSemi-supervised  learning
■653    ▼aTransfer  learning
■690    ▼a0308
■690    ▼a0489
■690    ▼a0769
■690    ▼a0715
■71020▼aHarvard  University▼bBiostatistics.
■7730  ▼tDissertations  Abstracts  International▼g86-12B.
■790    ▼a0084
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357664▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF18434 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.