서브메뉴
검색
Statistical Models for Alternative Splicing With Applications to Heterogeneous Disease
Statistical Models for Alternative Splicing With Applications to Heterogeneous Disease
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211151130
- ISBN
- 9798382830704
- DDC
- 574
- 저자명
- Wang, David.
- 서명/저자
- Statistical Models for Alternative Splicing With Applications to Heterogeneous Disease
- 발행사항
- [Sl] : University of Pennsylvania, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 203 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 85-12, Section: B.
- 주기사항
- Advisor: Barash, Yoseph.
- 학위논문주기
- Thesis (Ph.D.)--University of Pennsylvania, 2024.
- 초록/해제
- 요약This dissertation is divided into two distinct parts which are unified by a focus on developing novel statistical methods to analyze splicing data. In Chapter 2, we develop methods for subtype discovery in heterogeneous cancers. Identification of cancer subtypes characterized by actionable genetic lesions is a pivotal step for developing treatment and improving clinical care. However, in heterogeneous diseases such as Acute Myeloid Leukemia (AML), subtype discovery can be challenging since mutation burden, which has traditionally been prioritized for this task, is low. Recent studies pointing to splicing aberrations in AML motivate splicing based detection of cancer subtypes. We developed an unsupervised machine learning algorithm called CHESSBOARD to identify "tiles" defined by a subset of splicing events and patient samples that represent disease subtypes. The model allows for a flexible number of tiles, accounts for uncertainty of splicing quantification, and is able to model missing values as additional signals. We first apply CHESSBOARD to synthetic data to assess its domain specific modeling advantages, followed by analysis of several leukemia datasets. We show detected subtypes are reproducible in independent studies, investigate their possible regulatory drivers and probe their relation to known AML mutations. Finally, we demonstrate the potential clinical utility of CHESSBOARD by supplementing mutation based diagnostic assays with discovered splicing profiles to improve drug response correlation. In Chapter 3, we develop improved methods for discovery of splicing quantitative trait loci (sQTLs). Identification and characterization of sQTLs has emerged as a critical component in understanding the function of noncoding genetic variants implicated in disease. However, a significant number of sQTLs remain undiscovered due to limitations in both splicing quantification and statistical methods. Here we present a sQTL mapping framework that identifies thousands of novel variants that have been recurrently omitted in recent studies. Our method combines event and transcript level quantifications to identify variants associated with a more comprehensive set of splicing phenotypes, a regression model tailored for splicing data, and a hypothesis weighting method leveraging splicing specific covariates to improve sGene discovery power while controlling exact FWER. Using GTEX as a case study, we show that existing pipelines fail to report over 25% of sQTLs in comparison. We also introduce several techniques to improve downstream variant prioritization including multivariate fine mapping, effect size inference and visualization tools. Finally, in an application to GTEx data, we show that our pipeline discovers novel intron retention associated variants in the Alzheimer's CASS4 gene and variants in NAGNAG motifs. Furthermore, we show that newly discovered sQTLs co-localize with GWAS variants in the GWAS catalog for neurodegenerative disease. The newly discovered sQTL thus explains addition GWAS signal compared to existing approaches which provide novel insight into the functional role of genetic variants in splicing regulation.
- 일반주제명
- Bioinformatics
- 일반주제명
- Biostatistics
- 일반주제명
- Genetics
- 일반주제명
- Statistics
- 키워드
- Cancer subtypes
- 키워드
- Genetic variants
- 키워드
- Machine learning
- 키워드
- RNA splicing
- 기타저자
- University of Pennsylvania Genomics and Computational Biology
- 기본자료저록
- Dissertations Abstracts International. 85-12B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017160879
■00520250211151130
■006m o d
■007cr#unu||||||||
■020 ▼a9798382830704
■035 ▼a(MiAaPQ)AAI31147498
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a574
■1001 ▼aWang, David.
■24510▼aStatistical Models for Alternative Splicing With Applications to Heterogeneous Disease
■260 ▼a[Sl]▼bUniversity of Pennsylvania▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a203 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 85-12, Section: B.
■500 ▼aAdvisor: Barash, Yoseph.
■5021 ▼aThesis (Ph.D.)--University of Pennsylvania, 2024.
■520 ▼aThis dissertation is divided into two distinct parts which are unified by a focus on developing novel statistical methods to analyze splicing data. In Chapter 2, we develop methods for subtype discovery in heterogeneous cancers. Identification of cancer subtypes characterized by actionable genetic lesions is a pivotal step for developing treatment and improving clinical care. However, in heterogeneous diseases such as Acute Myeloid Leukemia (AML), subtype discovery can be challenging since mutation burden, which has traditionally been prioritized for this task, is low. Recent studies pointing to splicing aberrations in AML motivate splicing based detection of cancer subtypes. We developed an unsupervised machine learning algorithm called CHESSBOARD to identify "tiles" defined by a subset of splicing events and patient samples that represent disease subtypes. The model allows for a flexible number of tiles, accounts for uncertainty of splicing quantification, and is able to model missing values as additional signals. We first apply CHESSBOARD to synthetic data to assess its domain specific modeling advantages, followed by analysis of several leukemia datasets. We show detected subtypes are reproducible in independent studies, investigate their possible regulatory drivers and probe their relation to known AML mutations. Finally, we demonstrate the potential clinical utility of CHESSBOARD by supplementing mutation based diagnostic assays with discovered splicing profiles to improve drug response correlation. In Chapter 3, we develop improved methods for discovery of splicing quantitative trait loci (sQTLs). Identification and characterization of sQTLs has emerged as a critical component in understanding the function of noncoding genetic variants implicated in disease. However, a significant number of sQTLs remain undiscovered due to limitations in both splicing quantification and statistical methods. Here we present a sQTL mapping framework that identifies thousands of novel variants that have been recurrently omitted in recent studies. Our method combines event and transcript level quantifications to identify variants associated with a more comprehensive set of splicing phenotypes, a regression model tailored for splicing data, and a hypothesis weighting method leveraging splicing specific covariates to improve sGene discovery power while controlling exact FWER. Using GTEX as a case study, we show that existing pipelines fail to report over 25% of sQTLs in comparison. We also introduce several techniques to improve downstream variant prioritization including multivariate fine mapping, effect size inference and visualization tools. Finally, in an application to GTEx data, we show that our pipeline discovers novel intron retention associated variants in the Alzheimer's CASS4 gene and variants in NAGNAG motifs. Furthermore, we show that newly discovered sQTLs co-localize with GWAS variants in the GWAS catalog for neurodegenerative disease. The newly discovered sQTL thus explains addition GWAS signal compared to existing approaches which provide novel insight into the functional role of genetic variants in splicing regulation.
■590 ▼aSchool code: 0175.
■650 4▼aBioinformatics
■650 4▼aBiostatistics
■650 4▼aGenetics
■650 4▼aStatistics
■653 ▼aCancer subtypes
■653 ▼aGenetic variants
■653 ▼aMachine learning
■653 ▼aQuantitative trait loci
■653 ▼aRNA splicing
■653 ▼aStatistical methods
■690 ▼a0715
■690 ▼a0308
■690 ▼a0369
■690 ▼a0463
■71020▼aUniversity of Pennsylvania▼bGenomics and Computational Biology.
■7730 ▼tDissertations Abstracts International▼g85-12B.
■790 ▼a0175
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17160879▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


