서브메뉴
검색
Scalable Statistical Analysis and Improved Missing Data Imputation for High-Dimensional Omics Data
Scalable Statistical Analysis and Improved Missing Data Imputation for High-Dimensional Omics Data
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103617
- ISBN
- 9798291555965
- DDC
- 574
- 저자명
- Xia, Kai.
- 서명/저자
- Scalable Statistical Analysis and Improved Missing Data Imputation for High-Dimensional Omics Data
- 발행사항
- [Sl] : The University of North Carolina at Chapel Hill, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 118 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-02, Section: B.
- 주기사항
- Advisor: Zou, Fei;Zhao, Bingxin.
- 학위논문주기
- Thesis (Ph.D.)--The University of North Carolina at Chapel Hill, 2025.
- 초록/해제
- 요약Due to the special structure of large p and small n, with covariance structures between subjects in the current generation omics dataset, many statistical and computational algorithms face challenges of scalability in higher dimension.We develop novel approaches to solve three particular problems in high-dimensional settings: i) scale-up the computational efficiency of association analysis in twin studies; 2) improve the imputation performance for omics data by combining existing algorithm; 3) develop novel imputation algorithm via transfer-learning and fine-tuning approaches to address the block-missing problem in cross-omics cohort studies.In the first project, we address a common problem in twin studies that have been widely used in researches of inheritable diseases and traits. GWAS of multiple traits, such as eQTL studies in twins, requires association tests between thousands of transcripts and millions of single nucleotide polymorphisms (SNPs). Standard methods such as mixed-effects models are extremely computationally inefficient and impractical in the current generation of computers. We introduce TwinEQTL, a computationally efficient alternative to commonly used mixed effects models in the analysis of eQTL and GWAS in twin studies.In the second project, we move to missing data problem. Missing data is a common problem in various randomized and observational studies. Several approaches have been developed to handle missing data, and recently multiple imputation has become increasingly popular. Relying on the sophisticated associations among the variables in high-dimensional setting, we propose a novel two-phase approach to impute the missing entries by combining existing imputation algorithms. The proposed method can impute continuous data in high-dimensional space and has robust performance in different levels of missing rate.In the final project, we aim to address the persistent issue of block-missingness in omics data by leveraging pre-trained machine learning models for imputation, drawing on knowledge from large, existing multi-omics datasets. Our approach utilizes existing large-scale cross-omics studies and advanced machine learning algorithms to build robust and efficient models that can be pre-trained and fine-tuned for application to new datasets. Using transfer learning, we adapt and apply intricate patterns and dependencies learned from external data sources, enabling more accurate and context-aware imputation in target studies.
- 일반주제명
- Biostatistics
- 일반주제명
- Bioinformatics
- 일반주제명
- Genetics
- 키워드
- Omics data
- 기타저자
- The University of North Carolina at Chapel Hill Biostatistics
- 기본자료저록
- Dissertations Abstracts International. 87-02B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017357915
■00520260202103617
■006m o d
■007cr#unu||||||||
■020 ▼a9798291555965
■035 ▼a(MiAaPQ)AAI32044836
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a574
■1001 ▼aXia, Kai.
■24510▼aScalable Statistical Analysis and Improved Missing Data Imputation for High-Dimensional Omics Data
■260 ▼a[Sl]▼bThe University of North Carolina at Chapel Hill▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a118 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-02, Section: B.
■500 ▼aAdvisor: Zou, Fei;Zhao, Bingxin.
■5021 ▼aThesis (Ph.D.)--The University of North Carolina at Chapel Hill, 2025.
■520 ▼aDue to the special structure of large p and small n, with covariance structures between subjects in the current generation omics dataset, many statistical and computational algorithms face challenges of scalability in higher dimension.We develop novel approaches to solve three particular problems in high-dimensional settings: i) scale-up the computational efficiency of association analysis in twin studies; 2) improve the imputation performance for omics data by combining existing algorithm; 3) develop novel imputation algorithm via transfer-learning and fine-tuning approaches to address the block-missing problem in cross-omics cohort studies.In the first project, we address a common problem in twin studies that have been widely used in researches of inheritable diseases and traits. GWAS of multiple traits, such as eQTL studies in twins, requires association tests between thousands of transcripts and millions of single nucleotide polymorphisms (SNPs). Standard methods such as mixed-effects models are extremely computationally inefficient and impractical in the current generation of computers. We introduce TwinEQTL, a computationally efficient alternative to commonly used mixed effects models in the analysis of eQTL and GWAS in twin studies.In the second project, we move to missing data problem. Missing data is a common problem in various randomized and observational studies. Several approaches have been developed to handle missing data, and recently multiple imputation has become increasingly popular. Relying on the sophisticated associations among the variables in high-dimensional setting, we propose a novel two-phase approach to impute the missing entries by combining existing imputation algorithms. The proposed method can impute continuous data in high-dimensional space and has robust performance in different levels of missing rate.In the final project, we aim to address the persistent issue of block-missingness in omics data by leveraging pre-trained machine learning models for imputation, drawing on knowledge from large, existing multi-omics datasets. Our approach utilizes existing large-scale cross-omics studies and advanced machine learning algorithms to build robust and efficient models that can be pre-trained and fine-tuned for application to new datasets. Using transfer learning, we adapt and apply intricate patterns and dependencies learned from external data sources, enabling more accurate and context-aware imputation in target studies.
■590 ▼aSchool code: 0153.
■650 4▼aBiostatistics
■650 4▼aBioinformatics
■650 4▼aGenetics
■653 ▼aSingle nucleotide polymorphisms
■653 ▼aOmics data
■653 ▼aMissing data imputation
■653 ▼aHigh dimensional setting
■690 ▼a0308
■690 ▼a0800
■690 ▼a0369
■690 ▼a0715
■71020▼aThe University of North Carolina at Chapel Hill▼bBiostatistics.
■7730 ▼tDissertations Abstracts International▼g87-02B.
■790 ▼a0153
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357915▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


