본문

서브메뉴

Statistical and Computational Methods for High-Dimensional Genetics and Genomics Data
Statistical and Computational Methods for High-Dimensional Genetics and Genomics Data
Statistical and Computational Methods for High-Dimensional Genetics and Genomics Data

Detailed Information

자료유형  
 학위논문 서양
최종처리일시  
20260202105221
ISBN  
9798291566268
DDC  
574
저자명  
Wu, Peijun.
서명/저자  
Statistical and Computational Methods for High-Dimensional Genetics and Genomics Data
발행사항  
[Sl] : University of Michigan, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
275 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-03, Section: B.
주기사항  
Advisor: Zhou, Xiang.
학위논문주기  
Thesis (Ph.D.)--University of Michigan, 2025.
초록/해제  
요약Recent advances in transcriptomic technologies such as single-cell RNA sequencing (scRNA-seq), spatially resolved transcriptomics (SRT), and single-cell clustered regularly interspaced short palindromic repeats (sc-CRISPR) have enabled the measurement of gene expression. These technologies have opened doors for exploring gene functions, characterizing interactions between genes within biological pathways, and elucidating the molecular mechanisms underlying complex biological processes and diseases. In the meantime, they pose significant statistical and computational challenges for various analytic tasks, including the characterization of spatial transcriptomic landscapes within cell types, inference of causal gene regulatory networks, and expression quantitative trait loci (eQTL) analysis across cell types. In this dissertation, I develop several effective statistical methods to address these key analytical challenges, with the aim of advancing our understanding of gene functions in complex biological systems.In chapter II, I develop a new computational method, Celina, to identify cell type-specific spatially variable genes (ct-SVGs) which represent the genes that display diverse spatial expression patterns or are spatially variable within a specific cell type. Celina utilizes a spatially varying coefficient model to accurately capture each gene's spatial expression pattern in relation to the distribution of cell types across tissue locations, ensuring effective type I error control and high power. Celina proves powerful compared to existing methods in single-cell resolution spatial transcriptomics and stands as the only effective solution for spot-resolution spatial transcriptomics. Applied to five real datasets, Celina uncovers ct-SVGs associated with tumor progression and patient survival in lung cancer, identifies metagenes with unique spatial patterns linked to cell proliferation and immune response in kidney cancer, and detects genes preferentially expressed near amyloid-β plaques in an Alzheimer's model. The ct-SVGs detected by Celina open doors for novel biologically informed downstream analyses, unveiling functional cellular heterogeneity at an unprecedented scale.In chapter III, I first introduce an alternative framework for identifying downstream genes potentially causally influenced by perturbed target genes in sc-CRISPR studies. Specifically, I leverage the perturbation status of guide RNAs (gRNAs) in single cells as instrumental variables to establish a causal framework, enabling the identification of causal gene-gene relationships influenced by CRISPR-targeted genes through instrumental variable (IV) analysis. While the alternative framework improves type I error control, it is limited by challenges such as correlations within same cell population, weak gRNAs effects, and off-target effects, which can introduce bias, induce false positives, and compromise power. To overcome these limitations, I further propose Canon, a one-sample IV analysis method specifically tailored to systematically identify genes potentially causally influenced by perturbed target genes across diverse sc-CRISPR platforms. Canon ensures robust type I error control while maintaining high statistical power. I evaluated the performance of Canon through comprehensive simulations and applications to two real datasets. The gene-gene relationships identified by Canon provide valuable insights into the causal gene regulatory network, uncovering novel therapeutic targets for cancer treatment and demonstrating the transformative potential of sc-CRISPR screening to resolve causal networks at an unprecedented scale.In chapter IV, I introduce Joint-eQTL, a statistical framework that performs multi-cell-type eQTL mapping using pseudo-bulk data derived from scRNA-seq data, therefore enabling the joint inference of genotype effects across multiple cell types. Joint-eQTL employs a multivariate linear mixed model to capture both shared and cell-type-specific genotype effects across cell types. Importantly, Joint-eQTL contains three complementary statistical tests: (1) a genetic mean test that evaluates a shared genotype effect across cell types, (2) a genetic variance test that detects heterogeneous effects across cell types, and (3) a joint test that captures both shared and cell-type-specific genetic effects. Additionally, Joint-eQTL models the variation between individuals, leverages exact null distributions to ensure accurate type I error control, improves detection power and offers insights into the genetic basis of gene regulation across cellular contexts. We applied Joint-eQTL to two real datasets and identified functionally impactful genetic variants and novel biomarkers with reduced evolutionary conservation, thereby elucidating the underlying complex genetic architecture and regulatory mechanisms.
일반주제명  
Biostatistics
일반주제명  
Cellular biology
일반주제명  
Bioinformatics
일반주제명  
Genetics
키워드  
Statistical and computational methods
키워드  
Genomics
키워드  
Statistical methods
키워드  
Cell types
기타저자  
University of Michigan Biostatistics
기본자료저록  
Dissertations Abstracts International. 87-03B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017359832
■00520260202105221
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798291566268
■035    ▼a(MiAaPQ)AAI32271808
■035    ▼a(MiAaPQ)umichrackham006440
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a574
■1001  ▼aWu,  Peijun.
■24510▼aStatistical  and  Computational  Methods  for  High-Dimensional  Genetics  and  Genomics  Data
■260    ▼a[Sl]▼bUniversity  of  Michigan▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a275  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-03,  Section:  B.
■500    ▼aAdvisor:  Zhou,  Xiang.
■5021  ▼aThesis  (Ph.D.)--University  of  Michigan,  2025.
■520    ▼aRecent  advances  in  transcriptomic  technologies  such  as  single-cell  RNA  sequencing  (scRNA-seq),  spatially  resolved  transcriptomics  (SRT),  and  single-cell  clustered  regularly  interspaced  short  palindromic  repeats  (sc-CRISPR)  have  enabled  the  measurement  of  gene  expression.  These  technologies  have  opened  doors  for  exploring  gene  functions,  characterizing  interactions  between  genes  within  biological  pathways,  and  elucidating  the  molecular  mechanisms  underlying  complex  biological  processes  and  diseases.  In  the  meantime,  they  pose  significant  statistical  and  computational  challenges  for  various  analytic  tasks,  including  the  characterization  of  spatial  transcriptomic  landscapes  within  cell  types,  inference  of  causal  gene  regulatory  networks,  and  expression  quantitative  trait  loci  (eQTL)  analysis  across  cell  types.  In  this  dissertation,  I  develop  several  effective  statistical  methods  to  address  these  key  analytical  challenges,  with  the  aim  of  advancing  our  understanding  of  gene  functions  in  complex  biological  systems.In  chapter  II,  I  develop  a  new  computational  method,  Celina,  to  identify  cell  type-specific  spatially  variable  genes  (ct-SVGs)  which  represent  the  genes  that  display  diverse  spatial  expression  patterns  or  are  spatially  variable  within  a  specific  cell  type.  Celina  utilizes  a  spatially  varying  coefficient  model  to  accurately  capture  each  gene's  spatial  expression  pattern  in  relation  to  the  distribution  of  cell  types  across  tissue  locations,  ensuring  effective  type  I  error  control  and  high  power.  Celina  proves  powerful  compared  to  existing  methods  in  single-cell  resolution  spatial  transcriptomics  and  stands  as  the  only  effective  solution  for  spot-resolution  spatial  transcriptomics.  Applied  to  five  real  datasets,  Celina  uncovers  ct-SVGs  associated  with  tumor  progression  and  patient  survival  in  lung  cancer,  identifies  metagenes  with  unique  spatial  patterns  linked  to  cell  proliferation  and  immune  response  in  kidney  cancer,  and  detects  genes  preferentially  expressed  near  amyloid-β  plaques  in  an  Alzheimer's  model.  The  ct-SVGs  detected  by  Celina  open  doors  for  novel  biologically  informed  downstream  analyses,  unveiling  functional  cellular  heterogeneity  at  an  unprecedented  scale.In  chapter  III,  I  first  introduce  an  alternative  framework  for  identifying  downstream  genes  potentially  causally  influenced  by  perturbed  target  genes  in  sc-CRISPR  studies.  Specifically,  I  leverage  the  perturbation  status  of  guide  RNAs  (gRNAs)  in  single  cells  as  instrumental  variables  to  establish  a  causal  framework,  enabling  the  identification  of  causal  gene-gene  relationships  influenced  by  CRISPR-targeted  genes  through  instrumental  variable  (IV)  analysis.  While  the  alternative  framework  improves  type  I  error  control,  it  is  limited  by  challenges  such  as  correlations  within  same  cell  population,  weak  gRNAs  effects,  and  off-target  effects,  which  can  introduce  bias,  induce  false  positives,  and  compromise  power.  To  overcome  these  limitations,  I  further  propose  Canon,  a  one-sample  IV  analysis  method  specifically  tailored  to  systematically  identify  genes  potentially  causally  influenced  by  perturbed  target  genes  across  diverse  sc-CRISPR  platforms.  Canon  ensures  robust  type  I  error  control  while  maintaining  high  statistical  power.  I  evaluated  the  performance  of  Canon  through  comprehensive  simulations  and  applications  to  two  real  datasets.  The  gene-gene  relationships  identified  by  Canon  provide  valuable  insights  into  the  causal  gene  regulatory  network,  uncovering  novel  therapeutic  targets  for  cancer  treatment  and  demonstrating  the  transformative  potential  of  sc-CRISPR  screening  to  resolve  causal  networks  at  an  unprecedented  scale.In  chapter  IV,  I  introduce  Joint-eQTL,  a  statistical  framework  that  performs  multi-cell-type  eQTL  mapping  using  pseudo-bulk  data  derived  from  scRNA-seq  data,  therefore  enabling  the  joint  inference  of  genotype  effects  across  multiple  cell  types.  Joint-eQTL  employs  a  multivariate  linear  mixed  model  to  capture  both  shared  and  cell-type-specific  genotype  effects  across  cell  types.  Importantly,  Joint-eQTL  contains  three  complementary  statistical  tests:  (1)  a  genetic  mean  test  that  evaluates  a  shared  genotype  effect  across  cell  types,  (2)  a  genetic  variance  test  that  detects  heterogeneous  effects  across  cell  types,  and  (3)  a  joint  test  that  captures  both  shared  and  cell-type-specific  genetic  effects.  Additionally,  Joint-eQTL  models  the  variation  between  individuals,  leverages  exact  null  distributions  to  ensure  accurate  type  I  error  control,  improves  detection  power  and  offers  insights  into  the  genetic  basis  of  gene  regulation  across  cellular  contexts.  We  applied  Joint-eQTL  to  two  real  datasets  and  identified  functionally  impactful  genetic  variants  and  novel  biomarkers  with  reduced  evolutionary  conservation,  thereby  elucidating  the  underlying  complex  genetic  architecture  and  regulatory  mechanisms.
■590    ▼aSchool  code:  0127.
■650  4▼aBiostatistics
■650  4▼aCellular  biology
■650  4▼aBioinformatics
■650  4▼aGenetics
■653    ▼aStatistical  and  computational  methods
■653    ▼aGenomics
■653    ▼aStatistical  methods
■653    ▼aCell  types
■690    ▼a0308
■690    ▼a0379
■690    ▼a0369
■690    ▼a0715
■71020▼aUniversity  of  Michigan▼bBiostatistics.
■7730  ▼tDissertations  Abstracts  International▼g87-03B.
■790    ▼a0127
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17359832▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

Preview

Export

ChatGPT Discussion

AI Recommended Related Books


    New Books MORE
    Statistics for the past 3 years. Go to brief

    ค้นหาข้อมูลรายละเอียด

    • จองห้องพัก
    • ไม่อยู่
    • โฟลเดอร์ของฉัน
    • ขอดูแรก
    • Non-Book Loan Application
    • Nighttime Book Loan Application
    วัสดุ
    Reg No. Call No. ตำแหน่งที่ตั้ง สถานะ ยืมข้อมูล
    TF17915 전자도서 대출가능 My Folder 부재도서신고 비도서대출신청 야간 도서대출신청

    * จองมีอยู่ในหนังสือยืม เพื่อให้การสำรองที่นั่งคลิกที่ปุ่มจองห้องพัก

    Books borrowed together with this book

    Related Popular Books

    Available after logging in.