본문

서브메뉴

Dissecting the Function of the Non-Coding Genome Using Observational Data and Genomic Deep Learning Models
Dissecting the Function of the Non-Coding Genome Using Observational Data and Genomic Deep...
Dissecting the Function of the Non-Coding Genome Using Observational Data and Genomic Deep Learning Models

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202104825
ISBN  
9798293893133
DDC  
574
저자명  
Kathail, Pooja.
서명/저자  
Dissecting the Function of the Non-Coding Genome Using Observational Data and Genomic Deep Learning Models
발행사항  
[Sl] : University of California, Berkeley, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
111 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-04, Section: B.
주기사항  
Advisor: Ioannidis, Nilah;Ye, Jimmie.
학위논문주기  
Thesis (Ph.D.)--University of California, Berkeley, 2025.
초록/해제  
요약Understanding the causes of disease is a critical step to improving human health. This thesis shows how large-scale genomic datasets and machine learning methods can be used to understand the role of non-protein-coding genetic mutations in disease. In Chapter 2, we collect and generate a population-scale single cell multiome dataset from a diverse human cohort, and use this data to identify non-coding mutations associated with differences in gene expression (eQTLs) and chromatin accessibility (caQTLs). We find that caQTLs often explain more autoimmune disease signal than eQTLs, due to their ability to tag primed chromatin states. Next, we evaluate a recent machine learning paradigm---genomic deep learning---that models the relationship between non-coding sequences and cell type specific molecular phenotypes, such as gene expression and chromatin accessibility. A promising application of such models is in silico prediction of non-coding mutation effects. In Chapter 3, we find that current genomic deep learning models perform poorly in cell type specific regulatory elements, which harbor a large fraction of the heritability of complex diseases. We identify model training strategies---such as single-task learning---to maximize performance in cell type specific regulatory elements. Finally, in Chapter 4, we find that current genomic deep learning models have high uncertainty in their predictions for out of distribution sequences containing genetic variants, suggesting that variant-based training data may be a path towards improved variant effect prediction. Together, this work analyzes novel data measuring the effect of non-coding mutations on cell type specific gene regulation, providing insight into the data modalities most informative for disease, and exposes important limitations of current machine learning tools, furthering our ability to understand the role of non-coding mutations in disease.
일반주제명  
Bioinformatics
일반주제명  
Cellular biology
일반주제명  
Genetics
키워드  
Chromatin accessibility
키워드  
Genetic mutations
키워드  
Genetic variants
키워드  
Machine learning
키워드  
Gene regulation
기타저자  
University of California, Berkeley Bioinformatics & Computational Biology
기본자료저록  
Dissertations Abstracts International. 87-04B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017359041
■00520260202104825
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798293893133
■035    ▼a(MiAaPQ)AAI32169864
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a574
■1001  ▼aKathail,  Pooja.
■24510▼aDissecting  the  Function  of  the  Non-Coding  Genome  Using  Observational  Data  and  Genomic  Deep  Learning  Models
■260    ▼a[Sl]▼bUniversity  of  California,  Berkeley▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a111  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-04,  Section:  B.
■500    ▼aAdvisor:  Ioannidis,  Nilah;Ye,  Jimmie.
■5021  ▼aThesis  (Ph.D.)--University  of  California,  Berkeley,  2025.
■520    ▼aUnderstanding  the  causes  of  disease  is  a  critical  step  to  improving  human  health.  This  thesis  shows  how  large-scale  genomic  datasets  and  machine  learning  methods  can  be  used  to  understand  the  role  of  non-protein-coding  genetic  mutations  in  disease.  In  Chapter  2,  we  collect  and  generate  a  population-scale  single  cell  multiome  dataset  from  a  diverse  human  cohort,  and  use  this  data  to  identify  non-coding  mutations  associated  with  differences  in  gene  expression  (eQTLs)  and  chromatin  accessibility  (caQTLs).  We  find  that  caQTLs  often  explain  more  autoimmune  disease  signal  than  eQTLs,  due  to  their  ability  to  tag  primed  chromatin  states.  Next,  we  evaluate  a  recent  machine  learning  paradigm---genomic  deep  learning---that  models  the  relationship  between  non-coding  sequences  and  cell  type  specific  molecular  phenotypes,  such  as  gene  expression  and  chromatin  accessibility.  A  promising  application  of  such  models  is  in  silico  prediction  of  non-coding  mutation  effects.  In  Chapter  3,  we  find  that  current  genomic  deep  learning  models  perform  poorly  in  cell  type  specific  regulatory  elements,  which  harbor  a  large  fraction  of  the  heritability  of  complex  diseases.  We  identify  model  training  strategies---such  as  single-task  learning---to  maximize  performance  in  cell  type  specific  regulatory  elements.  Finally,  in  Chapter  4,  we  find  that  current  genomic  deep  learning  models  have  high  uncertainty  in  their  predictions  for  out  of  distribution  sequences  containing  genetic  variants,  suggesting  that  variant-based  training  data  may  be  a  path  towards  improved  variant  effect  prediction.  Together,  this  work  analyzes  novel  data  measuring  the  effect  of  non-coding  mutations  on  cell  type  specific  gene  regulation,  providing  insight  into  the  data  modalities  most  informative  for  disease,  and  exposes  important  limitations  of  current  machine  learning  tools,  furthering  our  ability  to  understand  the  role  of  non-coding  mutations  in  disease.
■590    ▼aSchool  code:  0028.
■650  4▼aBioinformatics
■650  4▼aCellular  biology
■650  4▼aGenetics
■653    ▼aChromatin  accessibility
■653    ▼aGenetic  mutations
■653    ▼aGenetic  variants
■653    ▼aMachine  learning
■653    ▼aGene  regulation
■690    ▼a0715
■690    ▼a0379
■690    ▼a0369
■71020▼aUniversity  of  California,  Berkeley▼bBioinformatics  &  Computational  Biology.
■7730  ▼tDissertations  Abstracts  International▼g87-04B.
■790    ▼a0028
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17359041▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF17280 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.