본문

서브메뉴

Powerful Statistical Methods for Precise Heritability Estimation and Partitioning Using Summary Statistics
Powerful Statistical Methods for Precise Heritability Estimation and Partitioning Using Su...
Powerful Statistical Methods for Precise Heritability Estimation and Partitioning Using Summary Statistics

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211151450
ISBN  
9798382784496
DDC  
575
저자명  
Li, Hui.
서명/저자  
Powerful Statistical Methods for Precise Heritability Estimation and Partitioning Using Summary Statistics
발행사항  
[Sl] : Harvard University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
241 p
주기사항  
Source: Dissertations Abstracts International, Volume: 85-12, Section: B.
주기사항  
Includes supplementary digital materials.
주기사항  
Advisor: Lin, Xihong.
학위논문주기  
Thesis (Ph.D.)--Harvard University, 2024.
초록/해제  
요약SNP-heritability estimation and partitioning are two of the most commonly performed analyses in statistical genetics. However, methods that use variant-level summary statistics for these estimation have low statistical efficiency, meaning that there is large uncertainty in the estimates produced from these methods. This uncertainty makes the results from downstream analyses less interpretable. Moreover, modeling and accounting for the linkage disequilibrium ("LD") or correlations between genetic markers in large-scale sequencing studies is difficult, as this correlation matrix can be expensive to store, share and compute with.This dissertation presents new methodological and algorithmic advances to address these challenges. Chapter I and II are dedicated to introducing likelihood-based approaches for efficient heritability estimation and heritability enrichment analyses, respectively. Chapter III is dedicated to an algorithmic development for low-dimensional representations of empirical correlation matrices from genomic studies, with the LD matrix approximation as a motivating example.In Chapter I, we introduce a new method for local heritability estimation -- Heritability Estimation with high Efficiency using LD and association Summary Statistics ("HEELS") - which significantly improves the statistical efficiency of summary-statistics-based heritability estimator. In a nutshell, the HEELS estimator is an iterative procedure based on transforming the well-established Henderson's algorithm for variance component estimation in linear mixed models (LMMs). It attains comparable statistical efficiency as the REML-based estimators which typically require access to individual-level data. In addition to introducing HEELS, we also propose a novel framework to approximate the empirical LD matrix using the sum of a low-rank matrix and a banded matrix. We show that this way of modeling the LD can reduce the cost of LD storage and effectively improve the computational efficiency of heritability estimation by HEELS.In Chapter II, we present "graphREML", a novel likelihood-based heritability enrichment estimator that operates on GWAS summary statistics and a sparse representation of the population LD matrix based on graphical models, allowing for overlapping and continuous annotations. The major method we compare graphREML against is stratified LD score regression ("S-LDSC"), a state-of-the-art method-of-moments estimator for heritability enrichment; graphREML improves upon S-LDSC by modeling the full likelihood of the summary statistics. To make our estimation procedure tractable and stable, we employ a second-order optimization method with an approximate Hessian and a trust-region algorithm. Compared to S-LDSC, graphREML is more powerful and identifies a larger number of significant enrichment (2.5 times more trait-annotation pairs). graphREML is applicable to summary association statistics for almost any trait, and its statistical efficiency will enable the identification of highly specific disease relevant functional features.In Chapter III, we build upon the LD approximation framework proposed in Chapter I, which represents the empirical LD as the sum of a banded and a low-rank matrix. We develop efficient algorithms to solve for the optimal Banded and Low-Rank representation of the Empirical ("BandaLoRE") correlation matrix, using coordinate descent and its variations. We found that BandaLoRE can significantly improve the computational efficiency of LD approximations and led to precise heritability estimates. We also explored the utility of our algorithms in other biological contexts, e.g., leveraging our approximation framework to model the contact matrices in Hi-C data analyses. We observed that BandaLoRE leads to highly efficient representations of the 3D interaction patterns on the genome. Although preliminary, this result showcases the potentially broader applicability and utility of our algorithm in both genetic and genomic studies.
일반주제명  
Genetics
일반주제명  
Statistics
일반주제명  
Biostatistics
키워드  
Genome-wide association studies
키워드  
Heritability estimation
키워드  
Linkage disequilibrium modeling
키워드  
Statistical genetics
기타저자  
Harvard University Biostatistics
기본자료저록  
Dissertations Abstracts International. 85-12B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017161822
■00520250211151450
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798382784496
■035    ▼a(MiAaPQ)AAI31296690
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a575
■1001  ▼aLi,  Hui.▼0(orcid)0000-0003-3625-2851
■24510▼aPowerful  Statistical  Methods  for  Precise  Heritability  Estimation  and  Partitioning  Using  Summary  Statistics
■260    ▼a[Sl]▼bHarvard  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a241  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  85-12,  Section:  B.
■500    ▼aIncludes  supplementary  digital  materials.
■500    ▼aAdvisor:  Lin,  Xihong.
■5021  ▼aThesis  (Ph.D.)--Harvard  University,  2024.
■520    ▼aSNP-heritability  estimation  and  partitioning  are  two  of  the  most  commonly  performed  analyses  in  statistical  genetics.  However,  methods  that  use  variant-level  summary  statistics  for  these  estimation  have  low  statistical  efficiency,  meaning  that  there  is  large  uncertainty  in  the  estimates  produced  from  these  methods.  This  uncertainty  makes  the  results  from  downstream  analyses  less  interpretable.  Moreover,  modeling  and  accounting  for  the  linkage  disequilibrium  ("LD")  or  correlations  between  genetic  markers  in  large-scale  sequencing  studies  is  difficult,  as  this  correlation  matrix  can  be  expensive  to  store,  share  and  compute  with.This  dissertation  presents  new  methodological  and  algorithmic  advances  to  address  these  challenges.  Chapter  I  and  II  are  dedicated  to  introducing  likelihood-based  approaches  for  efficient  heritability  estimation  and  heritability  enrichment  analyses,  respectively.  Chapter  III  is  dedicated  to  an  algorithmic  development  for  low-dimensional  representations  of  empirical  correlation  matrices  from  genomic  studies,  with  the  LD  matrix  approximation  as  a  motivating  example.In  Chapter  I,  we  introduce  a  new  method  for  local  heritability  estimation  --  Heritability  Estimation  with  high  Efficiency  using  LD  and  association  Summary  Statistics  ("HEELS")  -  which  significantly  improves  the  statistical  efficiency  of  summary-statistics-based  heritability  estimator.  In  a  nutshell,  the  HEELS  estimator  is  an  iterative  procedure  based  on  transforming  the  well-established  Henderson's  algorithm  for  variance  component  estimation  in  linear  mixed  models  (LMMs).  It  attains  comparable  statistical  efficiency  as  the  REML-based  estimators  which  typically  require  access  to  individual-level  data.  In  addition  to  introducing  HEELS,  we  also  propose  a  novel  framework  to  approximate  the  empirical  LD  matrix  using  the  sum  of  a  low-rank  matrix  and  a  banded  matrix.  We  show  that  this  way  of  modeling  the  LD  can  reduce  the  cost  of  LD  storage  and  effectively  improve  the  computational  efficiency  of  heritability  estimation  by  HEELS.In  Chapter  II,  we  present  "graphREML",  a  novel  likelihood-based  heritability  enrichment  estimator  that  operates  on  GWAS  summary  statistics  and  a  sparse  representation  of  the  population  LD  matrix  based  on  graphical  models,  allowing  for  overlapping  and  continuous  annotations.  The  major  method  we  compare  graphREML  against  is  stratified  LD  score  regression  ("S-LDSC"),  a  state-of-the-art  method-of-moments  estimator  for  heritability  enrichment;  graphREML  improves  upon  S-LDSC  by  modeling  the  full  likelihood  of  the  summary  statistics.  To  make  our  estimation  procedure  tractable  and  stable,  we  employ  a  second-order  optimization  method  with  an  approximate  Hessian  and  a  trust-region  algorithm.  Compared  to  S-LDSC,  graphREML  is  more  powerful  and  identifies  a  larger  number  of  significant  enrichment  (2.5  times  more  trait-annotation  pairs).  graphREML  is  applicable  to  summary  association  statistics  for  almost  any  trait,  and  its  statistical  efficiency  will  enable  the  identification  of  highly  specific  disease  relevant  functional  features.In  Chapter  III,  we  build  upon  the  LD  approximation  framework  proposed  in  Chapter  I,  which  represents  the  empirical  LD  as  the  sum  of  a  banded  and  a  low-rank  matrix.  We  develop  efficient  algorithms  to  solve  for  the  optimal  Banded  and  Low-Rank  representation  of  the  Empirical  ("BandaLoRE")  correlation  matrix,  using  coordinate  descent  and  its  variations.  We  found  that  BandaLoRE  can  significantly  improve  the  computational  efficiency  of  LD  approximations  and  led  to  precise  heritability  estimates.  We  also  explored  the  utility  of  our  algorithms  in  other  biological  contexts,  e.g.,  leveraging  our  approximation  framework  to  model  the  contact  matrices  in  Hi-C  data  analyses.  We  observed  that  BandaLoRE  leads  to  highly  efficient  representations  of  the  3D  interaction  patterns  on  the  genome.  Although  preliminary,  this  result  showcases  the  potentially  broader  applicability  and  utility  of  our  algorithm  in  both  genetic  and  genomic  studies.
■590    ▼aSchool  code:  0084.
■650  4▼aGenetics
■650  4▼aStatistics
■650  4▼aBiostatistics
■653    ▼aGenome-wide  association  studies
■653    ▼aHeritability  estimation
■653    ▼aLinkage  disequilibrium  modeling
■653    ▼aStatistical  genetics
■690    ▼a0369
■690    ▼a0463
■690    ▼a0308
■71020▼aHarvard  University▼bBiostatistics.
■7730  ▼tDissertations  Abstracts  International▼g85-12B.
■790    ▼a0084
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17161822▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF11510 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.