서브메뉴
검색
Ensemble Methods for Latent Structure Detection From Heterogeneous Genomic and Phenotypic Data
Ensemble Methods for Latent Structure Detection From Heterogeneous Genomic and Phenotypic Data
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202104758
- ISBN
- 9798265408518
- DDC
- 574
- 서명/저자
- Ensemble Methods for Latent Structure Detection From Heterogeneous Genomic and Phenotypic Data
- 발행사항
- [Sl] : Harvard University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 107 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-05, Section: B.
- 주기사항
- Advisor: Lin, Xihong.
- 학위논문주기
- Thesis (Ph.D.)--Harvard University, 2025.
- 초록/해제
- 요약Disentangling the hidden patterns within genomic and phenotypic data can improve our understanding of complex conditions. Recent methodological developments in statistics and machine learning have improved our ability to detect latent patterns in a variety of application areas; however, these methods are often unsuitable for some of the data types common to health and biomedical data. Likewise, many latent structure methods require prespecification of the dimensions of the latent space, which is typically unknown. In this work, we introduce three ensemble statistical and machine learning methods designed to fill in these gaps.In Chapter 1, we introduce LACE-UP (LAtent Class analysis Ensembled with Umap and Pca), an ensemble machine learning method that outperforms gold-standard and oracle methods for clustering multidimensional binary data. When applied to dietary behavior data from the UK Biobank, LACE-UP uncovers interpretable dietary subtypes that are associated with lipid levels and cardiovascular risk. In Chapter 2, we introduce SEEK-VEC (Spectral Ensembling of topic models with Eigenscore for K-agnostic Vocabulary Embedding and Classification), a spectral ensemble topic modeling method for count data that yields prioritization scores and grouping scores that enable variable classification, pattern detection, and model diagnostics. We show through simulations that SEEK-VEC outperforms standard methods, particularly in weaker signal strength settings. We apply SEEK-VEC to single-cell gene expression data, food preference questionnaire data, and self-reported psychopathology symptom data, and show that the method uncovers meaningful insights across a broad range of contexts. In Chapter 3, we introduce SEEK-VFI (Spectral Ensembling of topic models with Eigenscore for K-agnostic Variable Feature Identification), an extension of SEEK-VEC that ranks genes with respect to their relevance to cell trajectory structure. We show that SEEK-VFI outperforms leading methods for differentiating between trajectory-relevant and uninformative genes, and we apply SEEK-VFI to several single-cell RNA expression datasets and demonstrate its ability to recover the true trajectory structure within the data.This suite of methods, designed for non-continuous data, provide a lens into the latent structure underlying phenotypic and genomic data. These methods do not require the prespecification of the dimensions of the latent space and are robust to noise. Taken together, the promise of these methods and the development of similar methods in the future is a refined understanding of complex phenotypes and their underlying mechanisms, which in turn will improve diagnoses, prognoses, and care.
- 일반주제명
- Biostatistics
- 일반주제명
- Genetics
- 키워드
- Clustering
- 키워드
- Ensemble methods
- 키워드
- Topic modeling
- 기타저자
- Harvard University Biostatistics
- 기본자료저록
- Dissertations Abstracts International. 87-05B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017358831
■00520260202104758
■006m o d
■007cr#unu||||||||
■020 ▼a9798265408518
■035 ▼a(MiAaPQ)AAI32164108
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a574
■1001 ▼aDanning, Rebecca.
■24510▼aEnsemble Methods for Latent Structure Detection From Heterogeneous Genomic and Phenotypic Data
■260 ▼a[Sl]▼bHarvard University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a107 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-05, Section: B.
■500 ▼aAdvisor: Lin, Xihong.
■5021 ▼aThesis (Ph.D.)--Harvard University, 2025.
■520 ▼aDisentangling the hidden patterns within genomic and phenotypic data can improve our understanding of complex conditions. Recent methodological developments in statistics and machine learning have improved our ability to detect latent patterns in a variety of application areas; however, these methods are often unsuitable for some of the data types common to health and biomedical data. Likewise, many latent structure methods require prespecification of the dimensions of the latent space, which is typically unknown. In this work, we introduce three ensemble statistical and machine learning methods designed to fill in these gaps.In Chapter 1, we introduce LACE-UP (LAtent Class analysis Ensembled with Umap and Pca), an ensemble machine learning method that outperforms gold-standard and oracle methods for clustering multidimensional binary data. When applied to dietary behavior data from the UK Biobank, LACE-UP uncovers interpretable dietary subtypes that are associated with lipid levels and cardiovascular risk. In Chapter 2, we introduce SEEK-VEC (Spectral Ensembling of topic models with Eigenscore for K-agnostic Vocabulary Embedding and Classification), a spectral ensemble topic modeling method for count data that yields prioritization scores and grouping scores that enable variable classification, pattern detection, and model diagnostics. We show through simulations that SEEK-VEC outperforms standard methods, particularly in weaker signal strength settings. We apply SEEK-VEC to single-cell gene expression data, food preference questionnaire data, and self-reported psychopathology symptom data, and show that the method uncovers meaningful insights across a broad range of contexts. In Chapter 3, we introduce SEEK-VFI (Spectral Ensembling of topic models with Eigenscore for K-agnostic Variable Feature Identification), an extension of SEEK-VEC that ranks genes with respect to their relevance to cell trajectory structure. We show that SEEK-VFI outperforms leading methods for differentiating between trajectory-relevant and uninformative genes, and we apply SEEK-VFI to several single-cell RNA expression datasets and demonstrate its ability to recover the true trajectory structure within the data.This suite of methods, designed for non-continuous data, provide a lens into the latent structure underlying phenotypic and genomic data. These methods do not require the prespecification of the dimensions of the latent space and are robust to noise. Taken together, the promise of these methods and the development of similar methods in the future is a refined understanding of complex phenotypes and their underlying mechanisms, which in turn will improve diagnoses, prognoses, and care.
■590 ▼aSchool code: 0084.
■650 4▼aBiostatistics
■650 4▼aGenetics
■653 ▼aClustering
■653 ▼aDimension reduction
■653 ▼aEnsemble methods
■653 ▼aLatent variable analysis
■653 ▼aTopic modeling
■653 ▼aTrajectory analysis
■690 ▼a0308
■690 ▼a0369
■71020▼aHarvard University▼bBiostatistics.
■7730 ▼tDissertations Abstracts International▼g87-05B.
■790 ▼a0084
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17358831▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


