본문

서브메뉴

Distributed Algorithms and Statistical Inference for Multi-Site Analyses: Unfolding the Complexity of Heterogeneity in Real-World Data
Distributed Algorithms and Statistical Inference for Multi-Site Analyses: Unfolding the Co...
Distributed Algorithms and Statistical Inference for Multi-Site Analyses: Unfolding the Complexity of Heterogeneity in Real-World Data

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211151338
ISBN  
9798382830483
DDC  
574
저자명  
Tong, Jiayi.
서명/저자  
Distributed Algorithms and Statistical Inference for Multi-Site Analyses: Unfolding the Complexity of Heterogeneity in Real-World Data
발행사항  
[Sl] : University of Pennsylvania, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
147 p
주기사항  
Source: Dissertations Abstracts International, Volume: 85-12, Section: B.
주기사항  
Advisor: Chen, Yong.
학위논문주기  
Thesis (Ph.D.)--University of Pennsylvania, 2024.
초록/해제  
요약In the era of expanding real-world data (RWD) availability from distributed research networks (DRNs), leveraging large-scale data has become essential to generate evidence for clinical inquiries relevant to stakeholders within the healthcare system. To provide answers to questions about hospital and treatment options, medication queries, and others, we still face practical challenges in analyzing RWD, such as reporting bias, confounders, and rare events. It is particularly challenging to integrate data from multiple clinical sites within DRNs due to data privacy concerns, patient heterogeneity -- also known as the case-mix situation or patient-mix situation -- and the communication cost.In this work, centered on generating real-world clinical evidence from DRN data, our objective is to develop several distributed learning frameworks. These frameworks are specifically designed to provide insights into comparative effectiveness research, health system performance assessment, and evaluation of site-of-care-related racial disparities, all while addressing the complexities of patient heterogeneity in real-world data settings. In our initial study, we acknowledged the heterogeneity of event rates across multiple sites and proposed a distributed conditional logistic regression (dCLR) algorithm. By employing pairwise conditioning to eliminate site-specific parameters, this novel approach can account for heterogeneity between sites and lead to more robust estimations of regression coefficients.Advancing our research trajectory, we aim to develop an end-to-end framework that enhances the capability to perform specific downstream tasks. In our second body of work, we introduced the Distributed Hospital Comparer framework for EHR-based hospital profiling with distributed data, aiming to benefit the stakeholders (e.g., patients, providers, policymakers, and payers) in the healthcare systems. This framework consists of two major modules: a distributed learning module, namely dGEM (decentralized algorithm for the generalized linear mixed effects model), to address the dilemma of sharing individual patient-level data, and a counterfactual modeling module to tackle the case-mix variation of patients across different hospitals. The validity and applicability of this framework have been demonstrated using a centralized dataset from the U.S. Organ Procurement and Transplantation Network (OPTN) encompassing 149 centers. Subsequently, we applied the framework to a global study involving 12 sites across three countries within the OHDSI network, aiming to investigate variations in hospital performance, measured by COVID-19 mortality, across two pandemic periods.In the third part of our work, motivated by the existence of racial disparities in kidney transplant access and post-transplant outcomes between Non-Hispanic Black (NHB) and Non-Hispanic White (NHW) patients in the United States, we focused on studying the site of care, which is a key factor contributing to the racial disparities. In response, we developed a federated learning framework, named dGEM-disparity (decentralized algorithm for generalized linear mixed effect model for disparity evaluation) with the goal of assessing site-of-care-related racial disparities. This framework consists of two modules: the first module provides accurately estimated common effects and calibrated hospital-specific effects by requiring only aggregated data from each center, and the second adopts a counterfactual modeling approach to assess whether graft failure rates differ if NHB patients were admitted to transplant centers in the same distribution as NHW patients. This framework has been applied to the United States Renal Data System (USRDS) data from 39,043 adult patients across 73 transplant centers.
일반주제명  
Biostatistics
일반주제명  
Statistics
일반주제명  
Bioinformatics
키워드  
Data heterogeneity
키워드  
Distributed algorithms
키워드  
Multi-site analyses
키워드  
Real-world data
키워드  
Distributed research networks
기타저자  
University of Pennsylvania Epidemiology and Biostatistics
기본자료저록  
Dissertations Abstracts International. 85-12B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017161312
■00520250211151338
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798382830483
■035    ▼a(MiAaPQ)AAI31242038
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a574
■1001  ▼aTong,  Jiayi.
■24510▼aDistributed  Algorithms  and  Statistical  Inference  for  Multi-Site  Analyses:  Unfolding  the  Complexity  of  Heterogeneity  in  Real-World  Data
■260    ▼a[Sl]▼bUniversity  of  Pennsylvania▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a147  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  85-12,  Section:  B.
■500    ▼aAdvisor:  Chen,  Yong.
■5021  ▼aThesis  (Ph.D.)--University  of  Pennsylvania,  2024.
■520    ▼aIn  the  era  of  expanding  real-world  data  (RWD)  availability  from  distributed  research  networks  (DRNs),  leveraging  large-scale  data  has  become  essential  to  generate  evidence  for  clinical  inquiries  relevant  to  stakeholders  within  the  healthcare  system.  To  provide  answers  to  questions  about  hospital  and  treatment  options,  medication  queries,  and  others,  we  still  face  practical  challenges  in  analyzing  RWD,  such  as  reporting  bias,  confounders,  and  rare  events.  It  is  particularly  challenging  to  integrate  data  from  multiple  clinical  sites  within  DRNs  due  to  data  privacy  concerns,  patient  heterogeneity  --  also  known  as  the  case-mix  situation  or  patient-mix  situation  --  and  the  communication  cost.In  this  work,  centered  on  generating  real-world  clinical  evidence  from  DRN  data,  our  objective  is  to  develop  several  distributed  learning  frameworks.  These  frameworks  are  specifically  designed  to  provide  insights  into  comparative  effectiveness  research,  health  system  performance  assessment,  and  evaluation  of  site-of-care-related  racial  disparities,  all  while  addressing  the  complexities  of  patient  heterogeneity  in  real-world  data  settings.  In  our  initial  study,  we  acknowledged  the  heterogeneity  of  event  rates  across  multiple  sites  and  proposed  a  distributed  conditional  logistic  regression  (dCLR)  algorithm.  By  employing  pairwise  conditioning  to  eliminate  site-specific  parameters,  this  novel  approach  can  account  for  heterogeneity  between  sites  and  lead  to  more  robust  estimations  of  regression  coefficients.Advancing  our  research  trajectory,  we  aim  to  develop  an  end-to-end  framework  that  enhances  the  capability  to  perform  specific  downstream  tasks.  In  our  second  body  of  work,  we  introduced  the  Distributed  Hospital  Comparer  framework  for  EHR-based  hospital  profiling  with  distributed  data,  aiming  to  benefit  the  stakeholders  (e.g.,  patients,  providers,  policymakers,  and  payers)  in  the  healthcare  systems.  This  framework  consists  of  two  major  modules:  a  distributed  learning  module,  namely  dGEM  (decentralized  algorithm  for  the  generalized  linear  mixed  effects  model),  to  address  the  dilemma  of  sharing  individual  patient-level  data,  and  a  counterfactual  modeling  module  to  tackle  the  case-mix  variation  of  patients  across  different  hospitals.  The  validity  and  applicability  of  this  framework  have  been  demonstrated  using  a  centralized  dataset  from  the  U.S.  Organ  Procurement  and  Transplantation  Network  (OPTN)  encompassing  149  centers.  Subsequently,  we  applied  the  framework  to  a  global  study  involving  12  sites  across  three  countries  within  the  OHDSI  network,  aiming  to  investigate  variations  in  hospital  performance,  measured  by  COVID-19  mortality,  across  two  pandemic  periods.In  the  third  part  of  our  work,  motivated  by  the  existence  of  racial  disparities  in  kidney  transplant  access  and  post-transplant  outcomes  between  Non-Hispanic  Black  (NHB)  and  Non-Hispanic  White  (NHW)  patients  in  the  United  States,  we  focused  on  studying  the  site  of  care,  which  is  a  key  factor  contributing  to  the  racial  disparities.  In  response,  we  developed  a  federated  learning  framework,  named  dGEM-disparity  (decentralized  algorithm  for  generalized  linear  mixed  effect  model  for  disparity  evaluation)  with  the  goal  of  assessing  site-of-care-related  racial  disparities.  This  framework  consists  of  two  modules:  the  first  module  provides  accurately  estimated  common  effects  and  calibrated  hospital-specific  effects  by  requiring  only  aggregated  data  from  each  center,  and  the  second  adopts  a  counterfactual  modeling  approach  to  assess  whether  graft  failure  rates  differ  if  NHB  patients  were  admitted  to  transplant  centers  in  the  same  distribution  as  NHW  patients.  This  framework  has  been  applied  to  the  United  States  Renal  Data  System  (USRDS)  data  from  39,043  adult  patients  across  73  transplant  centers.
■590    ▼aSchool  code:  0175.
■650  4▼aBiostatistics
■650  4▼aStatistics
■650  4▼aBioinformatics
■653    ▼aData  heterogeneity
■653    ▼aDistributed  algorithms
■653    ▼aMulti-site  analyses
■653    ▼aReal-world  data
■653    ▼aDistributed  research  networks
■690    ▼a0308
■690    ▼a0769
■690    ▼a0715
■690    ▼a0463
■71020▼aUniversity  of  Pennsylvania▼bEpidemiology  and  Biostatistics.
■7730  ▼tDissertations  Abstracts  International▼g85-12B.
■790    ▼a0175
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17161312▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF12342 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.