본문

서브메뉴

Scalable Statistical Analysis and Improved Missing Data Imputation for High-Dimensional Omics Data
Scalable Statistical Analysis and Improved Missing Data Imputation for High-Dimensional Om...
Scalable Statistical Analysis and Improved Missing Data Imputation for High-Dimensional Omics Data

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202103617
ISBN  
9798291555965
DDC  
574
저자명  
Xia, Kai.
서명/저자  
Scalable Statistical Analysis and Improved Missing Data Imputation for High-Dimensional Omics Data
발행사항  
[Sl] : The University of North Carolina at Chapel Hill, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
118 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-02, Section: B.
주기사항  
Advisor: Zou, Fei;Zhao, Bingxin.
학위논문주기  
Thesis (Ph.D.)--The University of North Carolina at Chapel Hill, 2025.
초록/해제  
요약Due to the special structure of large p and small n, with covariance structures between subjects in the current generation omics dataset, many statistical and computational algorithms face challenges of scalability in higher dimension.We develop novel approaches to solve three particular problems in high-dimensional settings: i) scale-up the computational efficiency of association analysis in twin studies; 2) improve the imputation performance for omics data by combining existing algorithm; 3) develop novel imputation algorithm via transfer-learning and fine-tuning approaches to address the block-missing problem in cross-omics cohort studies.In the first project, we address a common problem in twin studies that have been widely used in researches of inheritable diseases and traits. GWAS of multiple traits, such as eQTL studies in twins, requires association tests between thousands of transcripts and millions of single nucleotide polymorphisms (SNPs). Standard methods such as mixed-effects models are extremely computationally inefficient and impractical in the current generation of computers. We introduce TwinEQTL, a computationally efficient alternative to commonly used mixed effects models in the analysis of eQTL and GWAS in twin studies.In the second project, we move to missing data problem. Missing data is a common problem in various randomized and observational studies. Several approaches have been developed to handle missing data, and recently multiple imputation has become increasingly popular. Relying on the sophisticated associations among the variables in high-dimensional setting, we propose a novel two-phase approach to impute the missing entries by combining existing imputation algorithms. The proposed method can impute continuous data in high-dimensional space and has robust performance in different levels of missing rate.In the final project, we aim to address the persistent issue of block-missingness in omics data by leveraging pre-trained machine learning models for imputation, drawing on knowledge from large, existing multi-omics datasets. Our approach utilizes existing large-scale cross-omics studies and advanced machine learning algorithms to build robust and efficient models that can be pre-trained and fine-tuned for application to new datasets. Using transfer learning, we adapt and apply intricate patterns and dependencies learned from external data sources, enabling more accurate and context-aware imputation in target studies.
일반주제명  
Biostatistics
일반주제명  
Bioinformatics
일반주제명  
Genetics
키워드  
Single nucleotide polymorphisms
키워드  
Omics data
키워드  
Missing data imputation
키워드  
High dimensional setting
기타저자  
The University of North Carolina at Chapel Hill Biostatistics
기본자료저록  
Dissertations Abstracts International. 87-02B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017357915
■00520260202103617
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798291555965
■035    ▼a(MiAaPQ)AAI32044836
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a574
■1001  ▼aXia,  Kai.
■24510▼aScalable  Statistical  Analysis  and  Improved  Missing  Data  Imputation  for  High-Dimensional  Omics  Data
■260    ▼a[Sl]▼bThe  University  of  North  Carolina  at  Chapel  Hill▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a118  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-02,  Section:  B.
■500    ▼aAdvisor:  Zou,  Fei;Zhao,  Bingxin.
■5021  ▼aThesis  (Ph.D.)--The  University  of  North  Carolina  at  Chapel  Hill,  2025.
■520    ▼aDue  to  the  special  structure  of  large  p  and  small  n,  with  covariance  structures  between  subjects  in  the  current  generation  omics  dataset,  many  statistical  and  computational  algorithms  face  challenges  of  scalability  in  higher  dimension.We  develop  novel  approaches  to  solve  three  particular  problems  in  high-dimensional  settings:  i)  scale-up  the  computational  efficiency  of  association  analysis  in  twin  studies;  2)  improve  the  imputation  performance  for  omics  data  by  combining  existing  algorithm;  3)  develop  novel  imputation  algorithm  via  transfer-learning  and  fine-tuning  approaches  to  address  the  block-missing  problem  in  cross-omics  cohort  studies.In  the  first  project,  we  address  a  common  problem  in  twin  studies  that  have  been  widely  used  in  researches  of  inheritable  diseases  and  traits.  GWAS  of  multiple  traits,  such  as  eQTL  studies  in  twins,  requires  association  tests  between  thousands  of  transcripts  and  millions  of  single  nucleotide  polymorphisms  (SNPs).  Standard  methods  such  as  mixed-effects  models  are  extremely  computationally  inefficient  and  impractical  in  the  current  generation  of  computers.  We  introduce  TwinEQTL,  a  computationally  efficient  alternative  to  commonly  used  mixed  effects  models  in  the  analysis  of  eQTL  and  GWAS  in  twin  studies.In  the  second  project,  we  move  to  missing  data  problem.  Missing  data  is  a  common  problem  in  various  randomized  and  observational  studies.  Several  approaches  have  been  developed  to  handle  missing  data,  and  recently  multiple  imputation  has  become  increasingly  popular.  Relying  on  the  sophisticated  associations  among  the  variables  in  high-dimensional  setting,  we  propose  a  novel  two-phase  approach  to  impute  the  missing  entries  by  combining  existing  imputation  algorithms.  The  proposed  method  can  impute  continuous  data  in  high-dimensional  space  and  has  robust  performance  in  different  levels  of  missing  rate.In  the  final  project,  we  aim  to  address  the  persistent  issue  of  block-missingness  in  omics  data  by  leveraging  pre-trained  machine  learning  models  for  imputation,  drawing  on  knowledge  from  large,  existing  multi-omics  datasets.  Our  approach  utilizes  existing  large-scale  cross-omics  studies  and  advanced  machine  learning  algorithms  to  build  robust  and  efficient  models  that  can  be  pre-trained  and  fine-tuned  for  application  to  new  datasets.  Using  transfer  learning,  we  adapt  and  apply  intricate  patterns  and  dependencies  learned  from  external  data  sources,  enabling  more  accurate  and  context-aware  imputation  in  target  studies.
■590    ▼aSchool  code:  0153.
■650  4▼aBiostatistics
■650  4▼aBioinformatics
■650  4▼aGenetics
■653    ▼aSingle  nucleotide  polymorphisms
■653    ▼aOmics  data
■653    ▼aMissing  data  imputation
■653    ▼aHigh  dimensional  setting
■690    ▼a0308
■690    ▼a0800
■690    ▼a0369
■690    ▼a0715
■71020▼aThe  University  of  North  Carolina  at  Chapel  Hill▼bBiostatistics.
■7730  ▼tDissertations  Abstracts  International▼g87-02B.
■790    ▼a0153
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357915▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF15191 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.