서브메뉴
검색
Causal Inference in Complex Observational Settings With Applications to Electronic Health Record Data
Causal Inference in Complex Observational Settings With Applications to Electronic Health Record Data
Detailed Information
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103543
- ISBN
- 9798280716261
- DDC
- 574
- 저자명
- Xu, Daniel.
- 서명/저자
- Causal Inference in Complex Observational Settings With Applications to Electronic Health Record Data
- 발행사항
- [Sl] : Harvard University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 208 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-12, Section: B.
- 주기사항
- Advisor: Mukherjee, Rajarshi.
- 학위논문주기
- Thesis (Ph.D.)--Harvard University, 2025.
- 초록/해제
- 요약Electronic health record (EHR) data serve as a valuable source of real-world evidence for assessing treatment effects, offering rich, longitudinal patient data that can enhance clinical research, inform healthcare decisions, and ultimately improve patient outcomes. However, leveraging these complex data sources for reliable statistical inference remains challenging. For example, one common difficulty is the limited availability of readily validated clinical outcomes in EHR data, which often requires labor-intensive manual annotation through chart review. A related challenge arises when studies span multiple institutions, where differences in patient populations can introduce heterogeneity and potential biases. These complications further exacerbate the fundamental challenge of using observational data to draw valid causal conclusions. In this dissertation, we propose novel statistical methods for causal inference that address certain complexities relevant to clinical studies leveraging EHR data. In particular, we focus on two settings of interest: (1) semi-supervised learning, where outcome labels are scarce, and (2) transfer learning, where data must be integrated across diverse populations. We hope that the contributions of this dissertation offer meaningful steps toward unlocking the full potential of EHR data for clinical and translational research.In Chapter 1, we address the problem of estimating treatment effects in a semi-supervised setting where labeled outcomes are sparse but unlabeled data are abundant. We develop a semi-supervised calibration method that leverages a subsample of labeled outcomes to calibrate inferred outcomes, ensuring that the downstream treatment effect estimator remains consistent despite potential errors in outcome imputation. Unlike most traditional semi-supervised methods, we allow the labeling mechanism to depend on the observed data, rather than assuming it is completely random. This problem is analogous to estimating mean outcomes in longitudinal studies with monotone missingness, and we show that our proposed estimator is asymptotically equivalent to the augmented inverse probability weighting (AIPW) estimator when a consistent estimate of the labeling propensity score is available. The estimator is multiply robust and locally semiparametric efficient. We also demonstrate improved finite-sample efficiency in semi-supervised settings, owing to an effective normalization of an implicit augmentation term. The finite-sample performance is evaluated through simulations, and we illustrate the method in a case study comparing the effectiveness of two anti-TNF therapies on remission outcomes in patients with rheumatoid arthritis.In Chapter 2, we consider the problem of causal mediation analysis in a semi-supervised setting. Causal mediation analysis is a fundamental tool for understanding how treatments or exposures affect outcomes through intermediate variables. However, existing methods lack theoretical and practical guarantees in the presence of substantial missingness in the outcome variable. To address this, we propose a robust and efficient method for the semi-supervised estimation of natural direct and indirect effects, accommodating settings both with and without surrogate outcomes. Our approach extends the double machine learning framework by constructing multiply robust one-step estimators, which enable the use of flexible machine learning algorithms for estimating nuisance functions. We show that the proposed estimators are asymptotically normal and locally minimax optimal under semi-supervised models that do not assume the distribution of the non-missing variables to be known. Through extensive simulations, we demonstrate that the estimators remain unbiased, yield valid inference, and achieve substantial efficiency gains by leveraging predictive surrogates, even under complex data-generating mechanisms. We illustrate the method using EHR data to examine whether racial disparities in Alzheimer's disease-related outcomes are mediated through cardiovascular comorbidities.In Chapter 3, we explore the problem of estimating low-dimensional functionals in a target population using data from both source and target domains, where the two distributions are only weakly aligned. This setting arises frequently in practice when performing transfer learning, yet remains underexplored in the context of functional estimation. We focus on three canonical functionals relevant to statistics and causal inference: the quadratic regression functional, the expected conditional covariance, and the mean response in missing data models. Rather than making strong parametric assumptions on the source and target distributions, we model weak alignment by assuming that the differences in conditional mean functions across domains are bounded in L2 norm. For each functional, we propose two estimation strategies: (1) an influence function-based estimator that employs existing minimax-optimal transfer learning methods for nuisance estimation, and (2) a functional-level confidence thresholding estimator that selects between source- and target-based estimators. We derive non-asymptotic upper bounds on the estimation risk in mean absolute error and show that both strategies achieve faster rates than target-only estimators, provided that posterior drift between domains is sufficiently small. When functionals involve two nuisances, we demonstrate that some of the proposed estimators are able to outperform the minimax-optimal target-only estimator as long as the degree of posterior drift in either of the two nuisance functions is sufficiently small, which we refer to as distributional shift double robustness. Furthermore, we demonstrate that under additional assumptions -- for example, by directly bounding the separation between the functional evaluated at source and target distributions -- the functional-level confidence thresholding estimator can attain minimax-optimal rates up to log factors.
- 일반주제명
- Biostatistics
- 일반주제명
- Bioinformatics
- 일반주제명
- Information technology
- 키워드
- Causal inference
- 기타저자
- Harvard University Biostatistics
- 기본자료저록
- Dissertations Abstracts International. 86-12B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017357664
■00520260202103543
■006m o d
■007cr#unu||||||||
■020 ▼a9798280716261
■035 ▼a(MiAaPQ)AAI32041071
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a574
■1001 ▼aXu, Daniel.▼0(orcid)0000-0002-4750-7138
■24510▼aCausal Inference in Complex Observational Settings With Applications to Electronic Health Record Data
■260 ▼a[Sl]▼bHarvard University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a208 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-12, Section: B.
■500 ▼aAdvisor: Mukherjee, Rajarshi.
■5021 ▼aThesis (Ph.D.)--Harvard University, 2025.
■520 ▼aElectronic health record (EHR) data serve as a valuable source of real-world evidence for assessing treatment effects, offering rich, longitudinal patient data that can enhance clinical research, inform healthcare decisions, and ultimately improve patient outcomes. However, leveraging these complex data sources for reliable statistical inference remains challenging. For example, one common difficulty is the limited availability of readily validated clinical outcomes in EHR data, which often requires labor-intensive manual annotation through chart review. A related challenge arises when studies span multiple institutions, where differences in patient populations can introduce heterogeneity and potential biases. These complications further exacerbate the fundamental challenge of using observational data to draw valid causal conclusions. In this dissertation, we propose novel statistical methods for causal inference that address certain complexities relevant to clinical studies leveraging EHR data. In particular, we focus on two settings of interest: (1) semi-supervised learning, where outcome labels are scarce, and (2) transfer learning, where data must be integrated across diverse populations. We hope that the contributions of this dissertation offer meaningful steps toward unlocking the full potential of EHR data for clinical and translational research.In Chapter 1, we address the problem of estimating treatment effects in a semi-supervised setting where labeled outcomes are sparse but unlabeled data are abundant. We develop a semi-supervised calibration method that leverages a subsample of labeled outcomes to calibrate inferred outcomes, ensuring that the downstream treatment effect estimator remains consistent despite potential errors in outcome imputation. Unlike most traditional semi-supervised methods, we allow the labeling mechanism to depend on the observed data, rather than assuming it is completely random. This problem is analogous to estimating mean outcomes in longitudinal studies with monotone missingness, and we show that our proposed estimator is asymptotically equivalent to the augmented inverse probability weighting (AIPW) estimator when a consistent estimate of the labeling propensity score is available. The estimator is multiply robust and locally semiparametric efficient. We also demonstrate improved finite-sample efficiency in semi-supervised settings, owing to an effective normalization of an implicit augmentation term. The finite-sample performance is evaluated through simulations, and we illustrate the method in a case study comparing the effectiveness of two anti-TNF therapies on remission outcomes in patients with rheumatoid arthritis.In Chapter 2, we consider the problem of causal mediation analysis in a semi-supervised setting. Causal mediation analysis is a fundamental tool for understanding how treatments or exposures affect outcomes through intermediate variables. However, existing methods lack theoretical and practical guarantees in the presence of substantial missingness in the outcome variable. To address this, we propose a robust and efficient method for the semi-supervised estimation of natural direct and indirect effects, accommodating settings both with and without surrogate outcomes. Our approach extends the double machine learning framework by constructing multiply robust one-step estimators, which enable the use of flexible machine learning algorithms for estimating nuisance functions. We show that the proposed estimators are asymptotically normal and locally minimax optimal under semi-supervised models that do not assume the distribution of the non-missing variables to be known. Through extensive simulations, we demonstrate that the estimators remain unbiased, yield valid inference, and achieve substantial efficiency gains by leveraging predictive surrogates, even under complex data-generating mechanisms. We illustrate the method using EHR data to examine whether racial disparities in Alzheimer's disease-related outcomes are mediated through cardiovascular comorbidities.In Chapter 3, we explore the problem of estimating low-dimensional functionals in a target population using data from both source and target domains, where the two distributions are only weakly aligned. This setting arises frequently in practice when performing transfer learning, yet remains underexplored in the context of functional estimation. We focus on three canonical functionals relevant to statistics and causal inference: the quadratic regression functional, the expected conditional covariance, and the mean response in missing data models. Rather than making strong parametric assumptions on the source and target distributions, we model weak alignment by assuming that the differences in conditional mean functions across domains are bounded in L2 norm. For each functional, we propose two estimation strategies: (1) an influence function-based estimator that employs existing minimax-optimal transfer learning methods for nuisance estimation, and (2) a functional-level confidence thresholding estimator that selects between source- and target-based estimators. We derive non-asymptotic upper bounds on the estimation risk in mean absolute error and show that both strategies achieve faster rates than target-only estimators, provided that posterior drift between domains is sufficiently small. When functionals involve two nuisances, we demonstrate that some of the proposed estimators are able to outperform the minimax-optimal target-only estimator as long as the degree of posterior drift in either of the two nuisance functions is sufficiently small, which we refer to as distributional shift double robustness. Furthermore, we demonstrate that under additional assumptions -- for example, by directly bounding the separation between the functional evaluated at source and target distributions -- the functional-level confidence thresholding estimator can attain minimax-optimal rates up to log factors.
■590 ▼aSchool code: 0084.
■650 4▼aBiostatistics
■650 4▼aBioinformatics
■650 4▼aInformation technology
■653 ▼aCausal inference
■653 ▼aElectronic health records
■653 ▼aSemi-supervised learning
■653 ▼aTransfer learning
■690 ▼a0308
■690 ▼a0489
■690 ▼a0769
■690 ▼a0715
■71020▼aHarvard University▼bBiostatistics.
■7730 ▼tDissertations Abstracts International▼g86-12B.
■790 ▼a0084
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357664▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.
Preview
Export
ChatGPT Discussion
AI Recommended Related Books
Подробнее информация.
- Бронирование
- не существует
- моя папка
- Первый запрос зрения
- Non-Book Loan Application
- Nighttime Book Loan Application
Available after logging in.


