본문

서브메뉴

Statistical Methods for Outcome Measurement Error Correction, and Multi-Study Prediction and Causal Inference Under Study Heterogeneity
Statistical Methods for Outcome Measurement Error Correction, and Multi-Study Prediction a...
Statistical Methods for Outcome Measurement Error Correction, and Multi-Study Prediction and Causal Inference Under Study Heterogeneity

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211152922
ISBN  
9798346532859
DDC  
574
저자명  
Wu, Yujie.
서명/저자  
Statistical Methods for Outcome Measurement Error Correction, and Multi-Study Prediction and Causal Inference Under Study Heterogeneity
발행사항  
[Sl] : Harvard University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
186 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-05, Section: B.
주기사항  
Advisor: Wang, Molin;Parmigiani, Giovanni.
학위논문주기  
Thesis (Ph.D.)--Harvard University, 2024.
초록/해제  
요약The availability of electronic health records (EHR) enables researchers to improve their understanding of disease etiology and prevention by identifying the health effect of risk factors and causal effects of treatment, and outcome predictions facilitated by machine learning tools can tailor health decision making to the individual level. However, EHR is often subject to measurement error and misclassification in exposures and outcomes, which will lead to biased estimates of health effects, with the effect typically being biased toward the null, causing false negative discoveries. On the other hand, accurate outcome prediction is often hindered by study-to-study variation that stems from different population studied, heterogeneous covariate-outcome relationships, causing a prediction model to have poor out-of-study prediction performance. Existing methods fail to address these issues, limiting researchers' ability to make use of the EHR data for disease prevention and treatment. This dissertation aims to develop statistical methods that aim to address the challenge in outcome measurement error for epidemiological studies and study heterogeneity for multi-study prediction and causal inference.In Chapter 1, we propose statistical methods based on weighted estimating equations to correct for outcome measurement errors-caused biases in effect estimates in the settings of time-to-event data with multiple failure types. We discuss the consistency and asymptotic normality of the proposed estimators and also derive their asymptotic variances. This work is motivated by the Conservation of Hearing Study which aims to evaluate risk factors for hearing loss in an ongoing cohort study, Nurses' Health Studies II. As an illustrative example, we apply the proposed method to adjust for the measurement errors in self- reported hearing outcomes when estimating the associations of tinnitus with hearing loss subtypes.In Chapter 2, we develop methods to analyze clustered competing risk data when event types are only available in a training dataset and are missing in the main study. We propose to estimate the exposure effects through the cause-specific proportional hazards frailty model, where random effects are introduced into the model to account for the within-cluster correlation. We propose a weighted penalized partial likelihood method where the weights represent the probabilities of the occurrence of events, and the weights can be obtained by fitting a classification model for the event types on the training dataset. Alternatively, we propose an imputation approach in which the missing event types are imputed based on the predictions from the classification model. We derive the analytical variances and evaluate the finite sample properties of our methods in an extensive simulation study. As an illustrative example, we apply our methods to estimate the associations between tinnitus and metabolic, sensory and metabolic+sensory hearing loss in the Conservation of Hearing Study Audiology Assessment Arm.In Chapter 3, we propose methods for multi-source domain adaptation (MSDA) for regression problems that leverage information from more than one source domain to make predictions in a target domain, where different domains may have different data distributions. First, we extend a flexible single-source DA algorithm for classification through outcome coarsening to enable its application to regression problems. We then augment our single-source DA algorithm for regression with ensemble learning to achieve multi-source DA. We consider three learning paradigms in the ensemble algorithm, which combines linearly the target-adapted learners trained with each source domain: (i) a multi-source stacking algorithm to obtain the ensemble weights; (ii) a similarity-based weighting where the weights reflect the quality of DA of each target-adapted learner; and (iii) a combination of the stacking and similarity weights. We illustrate the performance of our algorithms with simulations and a data application where the goal is to predict high-density lipoprotein (HDL) cholesterol levels using the gut microbiome. We observe a consistent improvement in prediction performance of our multi-source DA algorithm over the routinely used methods in all these scenarios.In Chapter 4, we propose methods to leverage information from multiple clinical trials to facilitate the estimation of study-specific causal effects.
일반주제명  
Biostatistics
키워드  
Electronic health records
키워드  
Health decision making
키워드  
Machine learning tools
기타저자  
Harvard University Biostatistics
기본자료저록  
Dissertations Abstracts International. 86-05B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017164203
■00520250211152922
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798346532859
■035    ▼a(MiAaPQ)AAI31559889
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a574
■1001  ▼aWu,  Yujie.▼0(orcid)0000-0002-2906-9998
■24510▼aStatistical  Methods  for  Outcome  Measurement  Error  Correction,  and  Multi-Study  Prediction  and  Causal  Inference  Under  Study  Heterogeneity
■260    ▼a[Sl]▼bHarvard  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a186  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-05,  Section:  B.
■500    ▼aAdvisor:  Wang,  Molin;Parmigiani,  Giovanni.
■5021  ▼aThesis  (Ph.D.)--Harvard  University,  2024.
■520    ▼aThe  availability  of  electronic  health  records  (EHR)  enables  researchers  to  improve  their  understanding  of  disease  etiology  and  prevention  by  identifying  the  health  effect  of  risk  factors  and  causal  effects  of  treatment,  and  outcome  predictions  facilitated  by  machine  learning  tools  can  tailor  health  decision  making  to  the  individual  level.  However,  EHR  is  often  subject  to  measurement  error  and  misclassification  in  exposures  and  outcomes,  which  will  lead  to  biased  estimates  of  health  effects,  with  the  effect  typically  being  biased  toward  the  null,  causing  false  negative  discoveries.  On  the  other  hand,  accurate  outcome  prediction  is  often  hindered  by  study-to-study  variation  that  stems  from  different  population  studied,  heterogeneous  covariate-outcome  relationships,  causing  a  prediction  model  to  have  poor  out-of-study  prediction  performance.  Existing  methods  fail  to  address  these  issues,  limiting  researchers'  ability  to  make  use  of  the  EHR  data  for  disease  prevention  and  treatment.  This  dissertation  aims  to  develop  statistical  methods  that  aim  to  address  the  challenge  in  outcome  measurement  error  for  epidemiological  studies  and  study  heterogeneity  for  multi-study  prediction  and  causal  inference.In  Chapter  1,  we  propose  statistical  methods  based  on  weighted  estimating  equations  to  correct  for  outcome  measurement  errors-caused  biases  in  effect  estimates  in  the  settings  of  time-to-event  data  with  multiple  failure  types.  We  discuss  the  consistency  and  asymptotic  normality  of  the  proposed  estimators  and  also  derive  their  asymptotic  variances.  This  work  is  motivated  by  the  Conservation  of  Hearing  Study  which  aims  to  evaluate  risk  factors  for  hearing  loss  in  an  ongoing  cohort  study,  Nurses'  Health  Studies  II.  As  an  illustrative  example,  we  apply  the  proposed  method  to  adjust  for  the  measurement  errors  in  self-  reported  hearing  outcomes  when  estimating  the  associations  of  tinnitus  with  hearing  loss  subtypes.In  Chapter  2,  we  develop  methods  to  analyze  clustered  competing  risk  data  when  event  types  are  only  available  in  a  training  dataset  and  are  missing  in  the  main  study.  We  propose  to  estimate  the  exposure  effects  through  the  cause-specific  proportional  hazards  frailty  model,  where  random  effects  are  introduced  into  the  model  to  account  for  the  within-cluster  correlation.  We  propose  a  weighted  penalized  partial  likelihood  method  where  the  weights  represent  the  probabilities  of  the  occurrence  of  events,  and  the  weights  can  be  obtained  by  fitting  a  classification  model  for  the  event  types  on  the  training  dataset.  Alternatively,  we  propose  an  imputation  approach  in  which  the  missing  event  types  are  imputed  based  on  the  predictions  from  the  classification  model.  We  derive  the  analytical  variances  and  evaluate  the  finite  sample  properties  of  our  methods  in  an  extensive  simulation  study.  As  an  illustrative  example,  we  apply  our  methods  to  estimate  the  associations  between  tinnitus  and  metabolic,  sensory  and  metabolic+sensory  hearing  loss  in  the  Conservation  of  Hearing  Study  Audiology  Assessment  Arm.In  Chapter  3,  we  propose  methods  for  multi-source  domain  adaptation  (MSDA)  for  regression  problems  that  leverage  information  from  more  than  one  source  domain  to  make  predictions  in  a  target  domain,  where  different  domains  may  have  different  data  distributions.  First,  we  extend  a  flexible  single-source  DA  algorithm  for  classification  through  outcome  coarsening  to  enable  its  application  to  regression  problems.  We  then  augment  our  single-source  DA  algorithm  for  regression  with  ensemble  learning  to  achieve  multi-source  DA.  We  consider  three  learning  paradigms  in  the  ensemble  algorithm,  which  combines  linearly  the  target-adapted  learners  trained  with  each  source  domain:  (i)  a  multi-source  stacking  algorithm  to  obtain  the  ensemble  weights;  (ii)  a  similarity-based  weighting  where  the  weights  reflect  the  quality  of  DA  of  each  target-adapted  learner;  and  (iii)  a  combination  of  the  stacking  and  similarity  weights.  We  illustrate  the  performance  of  our  algorithms  with  simulations  and  a  data  application  where  the  goal  is  to  predict  high-density  lipoprotein  (HDL)  cholesterol  levels  using  the  gut  microbiome.  We  observe  a  consistent  improvement  in  prediction  performance  of  our  multi-source  DA  algorithm  over  the  routinely  used  methods  in  all  these  scenarios.In  Chapter  4,  we  propose  methods  to  leverage  information  from  multiple  clinical  trials  to  facilitate  the  estimation  of  study-specific  causal  effects.
■590    ▼aSchool  code:  0084.
■650  4▼aBiostatistics
■653    ▼aElectronic  health  records
■653    ▼aHealth  decision  making
■653    ▼aMachine  learning  tools
■690    ▼a0308
■690    ▼a0769
■71020▼aHarvard  University▼bBiostatistics.
■7730  ▼tDissertations  Abstracts  International▼g86-05B.
■790    ▼a0084
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164203▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF10823 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.