본문

서브메뉴

Statistical Models for Alternative Splicing With Applications to Heterogeneous Disease
Statistical Models for Alternative Splicing With Applications to Heterogeneous Disease
Statistical Models for Alternative Splicing With Applications to Heterogeneous Disease

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211151130
ISBN  
9798382830704
DDC  
574
저자명  
Wang, David.
서명/저자  
Statistical Models for Alternative Splicing With Applications to Heterogeneous Disease
발행사항  
[Sl] : University of Pennsylvania, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
203 p
주기사항  
Source: Dissertations Abstracts International, Volume: 85-12, Section: B.
주기사항  
Advisor: Barash, Yoseph.
학위논문주기  
Thesis (Ph.D.)--University of Pennsylvania, 2024.
초록/해제  
요약This dissertation is divided into two distinct parts which are unified by a focus on developing novel statistical methods to analyze splicing data. In Chapter 2, we develop methods for subtype discovery in heterogeneous cancers. Identification of cancer subtypes characterized by actionable genetic lesions is a pivotal step for developing treatment and improving clinical care. However, in heterogeneous diseases such as Acute Myeloid Leukemia (AML), subtype discovery can be challenging since mutation burden, which has traditionally been prioritized for this task, is low. Recent studies pointing to splicing aberrations in AML motivate splicing based detection of cancer subtypes. We developed an unsupervised machine learning algorithm called CHESSBOARD to identify "tiles" defined by a subset of splicing events and patient samples that represent disease subtypes. The model allows for a flexible number of tiles, accounts for uncertainty of splicing quantification, and is able to model missing values as additional signals. We first apply CHESSBOARD to synthetic data to assess its domain specific modeling advantages, followed by analysis of several leukemia datasets. We show detected subtypes are reproducible in independent studies, investigate their possible regulatory drivers and probe their relation to known AML mutations. Finally, we demonstrate the potential clinical utility of CHESSBOARD by supplementing mutation based diagnostic assays with discovered splicing profiles to improve drug response correlation. In Chapter 3, we develop improved methods for discovery of splicing quantitative trait loci (sQTLs). Identification and characterization of sQTLs has emerged as a critical component in understanding the function of noncoding genetic variants implicated in disease. However, a significant number of sQTLs remain undiscovered due to limitations in both splicing quantification and statistical methods. Here we present a sQTL mapping framework that identifies thousands of novel variants that have been recurrently omitted in recent studies. Our method combines event and transcript level quantifications to identify variants associated with a more comprehensive set of splicing phenotypes, a regression model tailored for splicing data, and a hypothesis weighting method leveraging splicing specific covariates to improve sGene discovery power while controlling exact FWER. Using GTEX as a case study, we show that existing pipelines fail to report over 25% of sQTLs in comparison. We also introduce several techniques to improve downstream variant prioritization including multivariate fine mapping, effect size inference and visualization tools. Finally, in an application to GTEx data, we show that our pipeline discovers novel intron retention associated variants in the Alzheimer's CASS4 gene and variants in NAGNAG motifs. Furthermore, we show that newly discovered sQTLs co-localize with GWAS variants in the GWAS catalog for neurodegenerative disease. The newly discovered sQTL thus explains addition GWAS signal compared to existing approaches which provide novel insight into the functional role of genetic variants in splicing regulation. 
일반주제명  
Bioinformatics
일반주제명  
Biostatistics
일반주제명  
Genetics
일반주제명  
Statistics
키워드  
Cancer subtypes
키워드  
Genetic variants
키워드  
Machine learning
키워드  
Quantitative trait loci
키워드  
RNA splicing
키워드  
Statistical methods
기타저자  
University of Pennsylvania Genomics and Computational Biology
기본자료저록  
Dissertations Abstracts International. 85-12B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017160879
■00520250211151130
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798382830704
■035    ▼a(MiAaPQ)AAI31147498
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a574
■1001  ▼aWang,  David.
■24510▼aStatistical  Models  for  Alternative  Splicing  With  Applications  to  Heterogeneous  Disease
■260    ▼a[Sl]▼bUniversity  of  Pennsylvania▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a203  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  85-12,  Section:  B.
■500    ▼aAdvisor:  Barash,  Yoseph.
■5021  ▼aThesis  (Ph.D.)--University  of  Pennsylvania,  2024.
■520    ▼aThis  dissertation  is  divided  into  two  distinct  parts  which  are  unified  by  a  focus  on  developing  novel  statistical  methods  to  analyze  splicing  data.  In  Chapter  2,  we  develop  methods  for  subtype  discovery  in  heterogeneous  cancers.  Identification  of  cancer  subtypes  characterized  by  actionable  genetic  lesions  is  a  pivotal  step  for  developing  treatment  and  improving  clinical  care.  However,  in  heterogeneous  diseases  such  as  Acute  Myeloid  Leukemia  (AML),  subtype  discovery  can  be  challenging  since  mutation  burden,  which  has  traditionally  been  prioritized  for  this  task,  is  low.  Recent  studies  pointing  to  splicing  aberrations  in  AML  motivate  splicing  based  detection  of  cancer  subtypes.  We  developed  an  unsupervised  machine  learning  algorithm  called  CHESSBOARD  to  identify  "tiles"  defined  by  a  subset  of  splicing  events  and  patient  samples  that  represent  disease  subtypes.  The  model  allows  for  a  flexible  number  of  tiles,  accounts  for  uncertainty  of  splicing  quantification,  and  is  able  to  model  missing  values  as  additional  signals.  We  first  apply  CHESSBOARD  to  synthetic  data  to  assess  its  domain  specific  modeling  advantages,  followed  by  analysis  of  several  leukemia  datasets.  We  show  detected  subtypes  are  reproducible  in  independent  studies,  investigate  their  possible  regulatory  drivers  and  probe  their  relation  to  known  AML  mutations.  Finally,  we  demonstrate  the  potential  clinical  utility  of  CHESSBOARD  by  supplementing  mutation  based  diagnostic  assays  with  discovered  splicing  profiles  to  improve  drug  response  correlation.  In  Chapter  3,  we  develop  improved  methods  for  discovery  of  splicing  quantitative  trait  loci  (sQTLs).  Identification  and  characterization  of  sQTLs  has  emerged  as  a  critical  component  in  understanding  the  function  of  noncoding  genetic  variants  implicated  in  disease.  However,  a  significant  number  of  sQTLs  remain  undiscovered  due  to  limitations  in  both  splicing  quantification  and  statistical  methods.  Here  we  present  a  sQTL  mapping  framework  that  identifies  thousands  of  novel  variants  that  have  been  recurrently  omitted  in  recent  studies.  Our  method  combines  event  and  transcript  level  quantifications  to  identify  variants  associated  with  a  more  comprehensive  set  of  splicing  phenotypes,  a  regression  model  tailored  for  splicing  data,  and  a  hypothesis  weighting  method  leveraging  splicing  specific  covariates  to  improve  sGene  discovery  power  while  controlling  exact  FWER.  Using  GTEX  as  a  case  study,  we  show  that  existing  pipelines  fail  to  report  over  25%  of  sQTLs  in  comparison.  We  also  introduce  several  techniques  to  improve  downstream  variant  prioritization  including  multivariate  fine  mapping,  effect  size  inference  and  visualization  tools.  Finally,  in  an  application  to  GTEx  data,  we  show  that  our  pipeline  discovers  novel  intron  retention  associated  variants  in  the  Alzheimer's  CASS4  gene  and  variants  in  NAGNAG  motifs.  Furthermore,  we  show  that  newly  discovered  sQTLs  co-localize  with  GWAS  variants  in  the  GWAS  catalog  for  neurodegenerative  disease.  The  newly  discovered  sQTL  thus  explains  addition  GWAS  signal  compared  to  existing  approaches  which  provide  novel  insight  into  the  functional  role  of  genetic  variants  in  splicing  regulation. 
■590    ▼aSchool  code:  0175.
■650  4▼aBioinformatics
■650  4▼aBiostatistics
■650  4▼aGenetics
■650  4▼aStatistics
■653    ▼aCancer  subtypes
■653    ▼aGenetic  variants
■653    ▼aMachine  learning
■653    ▼aQuantitative  trait  loci
■653    ▼aRNA  splicing
■653    ▼aStatistical  methods
■690    ▼a0715
■690    ▼a0308
■690    ▼a0369
■690    ▼a0463
■71020▼aUniversity  of  Pennsylvania▼bGenomics  and  Computational  Biology.
■7730  ▼tDissertations  Abstracts  International▼g85-12B.
■790    ▼a0175
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17160879▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF13617 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.