서브메뉴
검색
Powerful Statistical Methods for Precise Heritability Estimation and Partitioning Using Summary Statistics
Powerful Statistical Methods for Precise Heritability Estimation and Partitioning Using Summary Statistics
Detailed Information
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211151450
- ISBN
- 9798382784496
- DDC
- 575
- 저자명
- Li, Hui.
- 서명/저자
- Powerful Statistical Methods for Precise Heritability Estimation and Partitioning Using Summary Statistics
- 발행사항
- [Sl] : Harvard University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 241 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 85-12, Section: B.
- 주기사항
- Includes supplementary digital materials.
- 주기사항
- Advisor: Lin, Xihong.
- 학위논문주기
- Thesis (Ph.D.)--Harvard University, 2024.
- 초록/해제
- 요약SNP-heritability estimation and partitioning are two of the most commonly performed analyses in statistical genetics. However, methods that use variant-level summary statistics for these estimation have low statistical efficiency, meaning that there is large uncertainty in the estimates produced from these methods. This uncertainty makes the results from downstream analyses less interpretable. Moreover, modeling and accounting for the linkage disequilibrium ("LD") or correlations between genetic markers in large-scale sequencing studies is difficult, as this correlation matrix can be expensive to store, share and compute with.This dissertation presents new methodological and algorithmic advances to address these challenges. Chapter I and II are dedicated to introducing likelihood-based approaches for efficient heritability estimation and heritability enrichment analyses, respectively. Chapter III is dedicated to an algorithmic development for low-dimensional representations of empirical correlation matrices from genomic studies, with the LD matrix approximation as a motivating example.In Chapter I, we introduce a new method for local heritability estimation -- Heritability Estimation with high Efficiency using LD and association Summary Statistics ("HEELS") - which significantly improves the statistical efficiency of summary-statistics-based heritability estimator. In a nutshell, the HEELS estimator is an iterative procedure based on transforming the well-established Henderson's algorithm for variance component estimation in linear mixed models (LMMs). It attains comparable statistical efficiency as the REML-based estimators which typically require access to individual-level data. In addition to introducing HEELS, we also propose a novel framework to approximate the empirical LD matrix using the sum of a low-rank matrix and a banded matrix. We show that this way of modeling the LD can reduce the cost of LD storage and effectively improve the computational efficiency of heritability estimation by HEELS.In Chapter II, we present "graphREML", a novel likelihood-based heritability enrichment estimator that operates on GWAS summary statistics and a sparse representation of the population LD matrix based on graphical models, allowing for overlapping and continuous annotations. The major method we compare graphREML against is stratified LD score regression ("S-LDSC"), a state-of-the-art method-of-moments estimator for heritability enrichment; graphREML improves upon S-LDSC by modeling the full likelihood of the summary statistics. To make our estimation procedure tractable and stable, we employ a second-order optimization method with an approximate Hessian and a trust-region algorithm. Compared to S-LDSC, graphREML is more powerful and identifies a larger number of significant enrichment (2.5 times more trait-annotation pairs). graphREML is applicable to summary association statistics for almost any trait, and its statistical efficiency will enable the identification of highly specific disease relevant functional features.In Chapter III, we build upon the LD approximation framework proposed in Chapter I, which represents the empirical LD as the sum of a banded and a low-rank matrix. We develop efficient algorithms to solve for the optimal Banded and Low-Rank representation of the Empirical ("BandaLoRE") correlation matrix, using coordinate descent and its variations. We found that BandaLoRE can significantly improve the computational efficiency of LD approximations and led to precise heritability estimates. We also explored the utility of our algorithms in other biological contexts, e.g., leveraging our approximation framework to model the contact matrices in Hi-C data analyses. We observed that BandaLoRE leads to highly efficient representations of the 3D interaction patterns on the genome. Although preliminary, this result showcases the potentially broader applicability and utility of our algorithm in both genetic and genomic studies.
- 일반주제명
- Genetics
- 일반주제명
- Statistics
- 일반주제명
- Biostatistics
- 기타저자
- Harvard University Biostatistics
- 기본자료저록
- Dissertations Abstracts International. 85-12B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017161822
■00520250211151450
■006m o d
■007cr#unu||||||||
■020 ▼a9798382784496
■035 ▼a(MiAaPQ)AAI31296690
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a575
■1001 ▼aLi, Hui.▼0(orcid)0000-0003-3625-2851
■24510▼aPowerful Statistical Methods for Precise Heritability Estimation and Partitioning Using Summary Statistics
■260 ▼a[Sl]▼bHarvard University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a241 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 85-12, Section: B.
■500 ▼aIncludes supplementary digital materials.
■500 ▼aAdvisor: Lin, Xihong.
■5021 ▼aThesis (Ph.D.)--Harvard University, 2024.
■520 ▼aSNP-heritability estimation and partitioning are two of the most commonly performed analyses in statistical genetics. However, methods that use variant-level summary statistics for these estimation have low statistical efficiency, meaning that there is large uncertainty in the estimates produced from these methods. This uncertainty makes the results from downstream analyses less interpretable. Moreover, modeling and accounting for the linkage disequilibrium ("LD") or correlations between genetic markers in large-scale sequencing studies is difficult, as this correlation matrix can be expensive to store, share and compute with.This dissertation presents new methodological and algorithmic advances to address these challenges. Chapter I and II are dedicated to introducing likelihood-based approaches for efficient heritability estimation and heritability enrichment analyses, respectively. Chapter III is dedicated to an algorithmic development for low-dimensional representations of empirical correlation matrices from genomic studies, with the LD matrix approximation as a motivating example.In Chapter I, we introduce a new method for local heritability estimation -- Heritability Estimation with high Efficiency using LD and association Summary Statistics ("HEELS") - which significantly improves the statistical efficiency of summary-statistics-based heritability estimator. In a nutshell, the HEELS estimator is an iterative procedure based on transforming the well-established Henderson's algorithm for variance component estimation in linear mixed models (LMMs). It attains comparable statistical efficiency as the REML-based estimators which typically require access to individual-level data. In addition to introducing HEELS, we also propose a novel framework to approximate the empirical LD matrix using the sum of a low-rank matrix and a banded matrix. We show that this way of modeling the LD can reduce the cost of LD storage and effectively improve the computational efficiency of heritability estimation by HEELS.In Chapter II, we present "graphREML", a novel likelihood-based heritability enrichment estimator that operates on GWAS summary statistics and a sparse representation of the population LD matrix based on graphical models, allowing for overlapping and continuous annotations. The major method we compare graphREML against is stratified LD score regression ("S-LDSC"), a state-of-the-art method-of-moments estimator for heritability enrichment; graphREML improves upon S-LDSC by modeling the full likelihood of the summary statistics. To make our estimation procedure tractable and stable, we employ a second-order optimization method with an approximate Hessian and a trust-region algorithm. Compared to S-LDSC, graphREML is more powerful and identifies a larger number of significant enrichment (2.5 times more trait-annotation pairs). graphREML is applicable to summary association statistics for almost any trait, and its statistical efficiency will enable the identification of highly specific disease relevant functional features.In Chapter III, we build upon the LD approximation framework proposed in Chapter I, which represents the empirical LD as the sum of a banded and a low-rank matrix. We develop efficient algorithms to solve for the optimal Banded and Low-Rank representation of the Empirical ("BandaLoRE") correlation matrix, using coordinate descent and its variations. We found that BandaLoRE can significantly improve the computational efficiency of LD approximations and led to precise heritability estimates. We also explored the utility of our algorithms in other biological contexts, e.g., leveraging our approximation framework to model the contact matrices in Hi-C data analyses. We observed that BandaLoRE leads to highly efficient representations of the 3D interaction patterns on the genome. Although preliminary, this result showcases the potentially broader applicability and utility of our algorithm in both genetic and genomic studies.
■590 ▼aSchool code: 0084.
■650 4▼aGenetics
■650 4▼aStatistics
■650 4▼aBiostatistics
■653 ▼aGenome-wide association studies
■653 ▼aHeritability estimation
■653 ▼aLinkage disequilibrium modeling
■653 ▼aStatistical genetics
■690 ▼a0369
■690 ▼a0463
■690 ▼a0308
■71020▼aHarvard University▼bBiostatistics.
■7730 ▼tDissertations Abstracts International▼g85-12B.
■790 ▼a0084
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17161822▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.
Preview
Export
ChatGPT Discussion
AI Recommended Related Books
Подробнее информация.
- Бронирование
- не существует
- моя папка
- Первый запрос зрения
- Non-Book Loan Application
- Nighttime Book Loan Application
Available after logging in.


