서브메뉴
검색
Modeling and Inference for Real-World Health Data: Addressing Complex Trajectories, Missingness, and Unstructured Text
Modeling and Inference for Real-World Health Data: Addressing Complex Trajectories, Missingness, and Unstructured Text
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105307
- ISBN
- 9798270289409
- DDC
- 574
- 저자명
- Qu, Yixiang.
- 서명/저자
- Modeling and Inference for Real-World Health Data: Addressing Complex Trajectories, Missingness, and Unstructured Text
- 발행사항
- [Sl] : The University of North Carolina at Chapel Hill, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 153 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-07, Section: B.
- 주기사항
- Advisor: Wu, Di;Ibrahim, Joseph G.
- 학위논문주기
- Thesis (Ph.D.)--The University of North Carolina at Chapel Hill, 2025.
- 초록/해제
- 요약The proliferation of complex, high-dimensional, and longitudinal data from sources like clinical trials and Electronic Health Records (EHRs) presents significant analytical challenges. Standard analytical methods are often insufficient, as their underlying assumptions fail to capture the nuanced dynamics of non-linear disease progression, complex microbial ecosystems, and unstructured clinical text. This dissertation addresses these limitations by developing and validating three novel computational frameworks, each tailored to the unique complexities of oncology clinical trials, longitudinal microbiome studies, and unstructured EHR analysis, respectively.The first project introduces a cure rate joint model for analyzing longitudinal tumor burden and time-to-event data in oncology. Addressing limitations in capturing non-linear treatment responses, this model integrates a random change-point structure to capture tumor shrinkage and regrowth, alongside a cure-rate component for long-term control. A primary innovation is leveraging survival data to constrain the change point's timing, ensuring regrowth precedes observed progression. Applied to a non-small cell lung cancer trial, this framework yields accurate treatment effect estimations and deeper insight into disease dynamics.The second project addresses irregularly-sampled data in longitudinal microbiome studies by proposing the Bidirectional GRU-ODE-Bayes (BGOB) model. This deep-learning framework uses Neural Ordinary Differential Equations (ODE) to learn continuous-time trajectories from sparse observations. By processing time-series bidirectionally, BGOB accurately interpolates missing time points and unobserved biological zeros to create complete data matrices. Applications to early childhood caries and inflammatory bowel disease cohorts demonstrate that BGOB significantly improves the statistical power of downstream analyses, such as differential abundance testing and clustering.The third project introduces precLLM, a framework enabling smaller, locally-deployed Large Language Models (LLMs) for unstructured EHR analysis to address privacy and cost barriers. The core innovation is a smart preprocessing step that filters long clinical notes to isolate relevant text before inference. Evaluations on private and public EHR datasets show that this preprocessing enhances the accuracy of smaller models for clinical phenotyping, often outperforming much larger models and proving more efficient than fine-tuning with limited data. We also discuss its potential for enabling more complex longitudinal analyses of patient trajectories within EHR data.
- 일반주제명
- Biostatistics
- 일반주제명
- Medicine
- 일반주제명
- Bioinformatics
- 일반주제명
- Clinical psychology
- 키워드
- EHR analysis
- 키워드
- Clinical trials
- 기타저자
- The University of North Carolina at Chapel Hill Biostatistics
- 기본자료저록
- Dissertations Abstracts International. 87-07B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017360118
■00520260202105307
■006m o d
■007cr#unu||||||||
■020 ▼a9798270289409
■035 ▼a(MiAaPQ)AAI32283323
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a574
■1001 ▼aQu, Yixiang.
■24510▼aModeling and Inference for Real-World Health Data: Addressing Complex Trajectories, Missingness, and Unstructured Text
■260 ▼a[Sl]▼bThe University of North Carolina at Chapel Hill▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a153 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-07, Section: B.
■500 ▼aAdvisor: Wu, Di;Ibrahim, Joseph G.
■5021 ▼aThesis (Ph.D.)--The University of North Carolina at Chapel Hill, 2025.
■520 ▼aThe proliferation of complex, high-dimensional, and longitudinal data from sources like clinical trials and Electronic Health Records (EHRs) presents significant analytical challenges. Standard analytical methods are often insufficient, as their underlying assumptions fail to capture the nuanced dynamics of non-linear disease progression, complex microbial ecosystems, and unstructured clinical text. This dissertation addresses these limitations by developing and validating three novel computational frameworks, each tailored to the unique complexities of oncology clinical trials, longitudinal microbiome studies, and unstructured EHR analysis, respectively.The first project introduces a cure rate joint model for analyzing longitudinal tumor burden and time-to-event data in oncology. Addressing limitations in capturing non-linear treatment responses, this model integrates a random change-point structure to capture tumor shrinkage and regrowth, alongside a cure-rate component for long-term control. A primary innovation is leveraging survival data to constrain the change point's timing, ensuring regrowth precedes observed progression. Applied to a non-small cell lung cancer trial, this framework yields accurate treatment effect estimations and deeper insight into disease dynamics.The second project addresses irregularly-sampled data in longitudinal microbiome studies by proposing the Bidirectional GRU-ODE-Bayes (BGOB) model. This deep-learning framework uses Neural Ordinary Differential Equations (ODE) to learn continuous-time trajectories from sparse observations. By processing time-series bidirectionally, BGOB accurately interpolates missing time points and unobserved biological zeros to create complete data matrices. Applications to early childhood caries and inflammatory bowel disease cohorts demonstrate that BGOB significantly improves the statistical power of downstream analyses, such as differential abundance testing and clustering.The third project introduces precLLM, a framework enabling smaller, locally-deployed Large Language Models (LLMs) for unstructured EHR analysis to address privacy and cost barriers. The core innovation is a smart preprocessing step that filters long clinical notes to isolate relevant text before inference. Evaluations on private and public EHR datasets show that this preprocessing enhances the accuracy of smaller models for clinical phenotyping, often outperforming much larger models and proving more efficient than fine-tuning with limited data. We also discuss its potential for enabling more complex longitudinal analyses of patient trajectories within EHR data.
■590 ▼aSchool code: 0153.
■650 4▼aBiostatistics
■650 4▼aMedicine
■650 4▼aBioinformatics
■650 4▼aClinical psychology
■653 ▼aElectronic Health Records
■653 ▼aEHR analysis
■653 ▼aClinical trials
■653 ▼aLongitudinal microbiome studies
■690 ▼a0308
■690 ▼a0564
■690 ▼a0622
■690 ▼a0715
■71020▼aThe University of North Carolina at Chapel Hill▼bBiostatistics.
■7730 ▼tDissertations Abstracts International▼g87-07B.
■790 ▼a0153
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17360118▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


