서브메뉴
검색
Statistical and Computational Methods for High-Dimensional Genetics and Genomics Data
Statistical and Computational Methods for High-Dimensional Genetics and Genomics Data
Detailed Information
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105221
- ISBN
- 9798291566268
- DDC
- 574
- 저자명
- Wu, Peijun.
- 서명/저자
- Statistical and Computational Methods for High-Dimensional Genetics and Genomics Data
- 발행사항
- [Sl] : University of Michigan, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 275 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-03, Section: B.
- 주기사항
- Advisor: Zhou, Xiang.
- 학위논문주기
- Thesis (Ph.D.)--University of Michigan, 2025.
- 초록/해제
- 요약Recent advances in transcriptomic technologies such as single-cell RNA sequencing (scRNA-seq), spatially resolved transcriptomics (SRT), and single-cell clustered regularly interspaced short palindromic repeats (sc-CRISPR) have enabled the measurement of gene expression. These technologies have opened doors for exploring gene functions, characterizing interactions between genes within biological pathways, and elucidating the molecular mechanisms underlying complex biological processes and diseases. In the meantime, they pose significant statistical and computational challenges for various analytic tasks, including the characterization of spatial transcriptomic landscapes within cell types, inference of causal gene regulatory networks, and expression quantitative trait loci (eQTL) analysis across cell types. In this dissertation, I develop several effective statistical methods to address these key analytical challenges, with the aim of advancing our understanding of gene functions in complex biological systems.In chapter II, I develop a new computational method, Celina, to identify cell type-specific spatially variable genes (ct-SVGs) which represent the genes that display diverse spatial expression patterns or are spatially variable within a specific cell type. Celina utilizes a spatially varying coefficient model to accurately capture each gene's spatial expression pattern in relation to the distribution of cell types across tissue locations, ensuring effective type I error control and high power. Celina proves powerful compared to existing methods in single-cell resolution spatial transcriptomics and stands as the only effective solution for spot-resolution spatial transcriptomics. Applied to five real datasets, Celina uncovers ct-SVGs associated with tumor progression and patient survival in lung cancer, identifies metagenes with unique spatial patterns linked to cell proliferation and immune response in kidney cancer, and detects genes preferentially expressed near amyloid-β plaques in an Alzheimer's model. The ct-SVGs detected by Celina open doors for novel biologically informed downstream analyses, unveiling functional cellular heterogeneity at an unprecedented scale.In chapter III, I first introduce an alternative framework for identifying downstream genes potentially causally influenced by perturbed target genes in sc-CRISPR studies. Specifically, I leverage the perturbation status of guide RNAs (gRNAs) in single cells as instrumental variables to establish a causal framework, enabling the identification of causal gene-gene relationships influenced by CRISPR-targeted genes through instrumental variable (IV) analysis. While the alternative framework improves type I error control, it is limited by challenges such as correlations within same cell population, weak gRNAs effects, and off-target effects, which can introduce bias, induce false positives, and compromise power. To overcome these limitations, I further propose Canon, a one-sample IV analysis method specifically tailored to systematically identify genes potentially causally influenced by perturbed target genes across diverse sc-CRISPR platforms. Canon ensures robust type I error control while maintaining high statistical power. I evaluated the performance of Canon through comprehensive simulations and applications to two real datasets. The gene-gene relationships identified by Canon provide valuable insights into the causal gene regulatory network, uncovering novel therapeutic targets for cancer treatment and demonstrating the transformative potential of sc-CRISPR screening to resolve causal networks at an unprecedented scale.In chapter IV, I introduce Joint-eQTL, a statistical framework that performs multi-cell-type eQTL mapping using pseudo-bulk data derived from scRNA-seq data, therefore enabling the joint inference of genotype effects across multiple cell types. Joint-eQTL employs a multivariate linear mixed model to capture both shared and cell-type-specific genotype effects across cell types. Importantly, Joint-eQTL contains three complementary statistical tests: (1) a genetic mean test that evaluates a shared genotype effect across cell types, (2) a genetic variance test that detects heterogeneous effects across cell types, and (3) a joint test that captures both shared and cell-type-specific genetic effects. Additionally, Joint-eQTL models the variation between individuals, leverages exact null distributions to ensure accurate type I error control, improves detection power and offers insights into the genetic basis of gene regulation across cellular contexts. We applied Joint-eQTL to two real datasets and identified functionally impactful genetic variants and novel biomarkers with reduced evolutionary conservation, thereby elucidating the underlying complex genetic architecture and regulatory mechanisms.
- 일반주제명
- Biostatistics
- 일반주제명
- Cellular biology
- 일반주제명
- Bioinformatics
- 일반주제명
- Genetics
- 키워드
- Genomics
- 키워드
- Cell types
- 기타저자
- University of Michigan Biostatistics
- 기본자료저록
- Dissertations Abstracts International. 87-03B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017359832
■00520260202105221
■006m o d
■007cr#unu||||||||
■020 ▼a9798291566268
■035 ▼a(MiAaPQ)AAI32271808
■035 ▼a(MiAaPQ)umichrackham006440
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a574
■1001 ▼aWu, Peijun.
■24510▼aStatistical and Computational Methods for High-Dimensional Genetics and Genomics Data
■260 ▼a[Sl]▼bUniversity of Michigan▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a275 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-03, Section: B.
■500 ▼aAdvisor: Zhou, Xiang.
■5021 ▼aThesis (Ph.D.)--University of Michigan, 2025.
■520 ▼aRecent advances in transcriptomic technologies such as single-cell RNA sequencing (scRNA-seq), spatially resolved transcriptomics (SRT), and single-cell clustered regularly interspaced short palindromic repeats (sc-CRISPR) have enabled the measurement of gene expression. These technologies have opened doors for exploring gene functions, characterizing interactions between genes within biological pathways, and elucidating the molecular mechanisms underlying complex biological processes and diseases. In the meantime, they pose significant statistical and computational challenges for various analytic tasks, including the characterization of spatial transcriptomic landscapes within cell types, inference of causal gene regulatory networks, and expression quantitative trait loci (eQTL) analysis across cell types. In this dissertation, I develop several effective statistical methods to address these key analytical challenges, with the aim of advancing our understanding of gene functions in complex biological systems.In chapter II, I develop a new computational method, Celina, to identify cell type-specific spatially variable genes (ct-SVGs) which represent the genes that display diverse spatial expression patterns or are spatially variable within a specific cell type. Celina utilizes a spatially varying coefficient model to accurately capture each gene's spatial expression pattern in relation to the distribution of cell types across tissue locations, ensuring effective type I error control and high power. Celina proves powerful compared to existing methods in single-cell resolution spatial transcriptomics and stands as the only effective solution for spot-resolution spatial transcriptomics. Applied to five real datasets, Celina uncovers ct-SVGs associated with tumor progression and patient survival in lung cancer, identifies metagenes with unique spatial patterns linked to cell proliferation and immune response in kidney cancer, and detects genes preferentially expressed near amyloid-β plaques in an Alzheimer's model. The ct-SVGs detected by Celina open doors for novel biologically informed downstream analyses, unveiling functional cellular heterogeneity at an unprecedented scale.In chapter III, I first introduce an alternative framework for identifying downstream genes potentially causally influenced by perturbed target genes in sc-CRISPR studies. Specifically, I leverage the perturbation status of guide RNAs (gRNAs) in single cells as instrumental variables to establish a causal framework, enabling the identification of causal gene-gene relationships influenced by CRISPR-targeted genes through instrumental variable (IV) analysis. While the alternative framework improves type I error control, it is limited by challenges such as correlations within same cell population, weak gRNAs effects, and off-target effects, which can introduce bias, induce false positives, and compromise power. To overcome these limitations, I further propose Canon, a one-sample IV analysis method specifically tailored to systematically identify genes potentially causally influenced by perturbed target genes across diverse sc-CRISPR platforms. Canon ensures robust type I error control while maintaining high statistical power. I evaluated the performance of Canon through comprehensive simulations and applications to two real datasets. The gene-gene relationships identified by Canon provide valuable insights into the causal gene regulatory network, uncovering novel therapeutic targets for cancer treatment and demonstrating the transformative potential of sc-CRISPR screening to resolve causal networks at an unprecedented scale.In chapter IV, I introduce Joint-eQTL, a statistical framework that performs multi-cell-type eQTL mapping using pseudo-bulk data derived from scRNA-seq data, therefore enabling the joint inference of genotype effects across multiple cell types. Joint-eQTL employs a multivariate linear mixed model to capture both shared and cell-type-specific genotype effects across cell types. Importantly, Joint-eQTL contains three complementary statistical tests: (1) a genetic mean test that evaluates a shared genotype effect across cell types, (2) a genetic variance test that detects heterogeneous effects across cell types, and (3) a joint test that captures both shared and cell-type-specific genetic effects. Additionally, Joint-eQTL models the variation between individuals, leverages exact null distributions to ensure accurate type I error control, improves detection power and offers insights into the genetic basis of gene regulation across cellular contexts. We applied Joint-eQTL to two real datasets and identified functionally impactful genetic variants and novel biomarkers with reduced evolutionary conservation, thereby elucidating the underlying complex genetic architecture and regulatory mechanisms.
■590 ▼aSchool code: 0127.
■650 4▼aBiostatistics
■650 4▼aCellular biology
■650 4▼aBioinformatics
■650 4▼aGenetics
■653 ▼aStatistical and computational methods
■653 ▼aGenomics
■653 ▼aStatistical methods
■653 ▼aCell types
■690 ▼a0308
■690 ▼a0379
■690 ▼a0369
■690 ▼a0715
■71020▼aUniversity of Michigan▼bBiostatistics.
■7730 ▼tDissertations Abstracts International▼g87-03B.
■790 ▼a0127
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17359832▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.
Preview
Export
ChatGPT Discussion
AI Recommended Related Books
ค้นหาข้อมูลรายละเอียด
- จองห้องพัก
- ไม่อยู่
- โฟลเดอร์ของฉัน
- ขอดูแรก
- Non-Book Loan Application
- Nighttime Book Loan Application
Available after logging in.


