본문

서브메뉴

Improving Select Applications of Long-Read DNA Sequencing
Improving Select Applications of Long-Read DNA Sequencing
Improving Select Applications of Long-Read DNA Sequencing

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211152952
ISBN  
9798384042013
DDC  
004
저자명  
Dunn, Timothy J.
서명/저자  
Improving Select Applications of Long-Read DNA Sequencing
발행사항  
[Sl] : University of Michigan, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
184 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-03, Section: B.
주기사항  
Advisor: Narayanasamy, Satish.
학위논문주기  
Thesis (Ph.D.)--University of Michigan, 2024.
초록/해제  
요약The cost to sequence a human genome has dropped from an estimated $300 million to well under $1,000 over the past two decades. In fact, several companies - which can amortize sequencing cost by running hundreds of samples in parallel - have recently claimed to have reached the $100 human genome. Following such dramatic cost improvements, whole genome sequencing is just starting to regularly be used for cancer profiling, rare genetic disease detection, agricultural breeding, pathogen detection, microbiome bacterial abundance estimation, evolutionary biology research, personalized medicine development, and much more. As sequencing costs continue to drop, as accuracy improves, and as new applications are discovered, DNA sequencing will become increasingly ubiquitous. At the same time that sequencing cost is rapidly dropping, new sequencing technologies have emerged that offer greater capabilities than ever before. In particular, nanopore-based long-read sequencing has no theoretical limit on the length of a contiguous DNA sequence, or "read", that can be measured. In comparison to short-read sequencing technologies that have dominated the sequencing market thus far (with maximum read lengths of 100 base pairs), nanopore devices have sequenced entire bacterial chromosomes in a single strand. The current nanopore read length record stands at over four million bases. Longer read lengths result in fewer problems during read mapping and genome assembly, allowing insight into complex regions of the genome and types of genetic variation that have been historically under-studied. Although nanopore devices were originally limited by their approximately 80% per-base accuracy when first publicly released in 2015, this accuracy has increased to over 99% in recent years with the adoption of deep learning basecallers. Nanopore-based sequencing devices are also the first to come in a portable handheld form factor and offer real-time analysis of raw data as it is being recorded, further expanding potential use cases. Despite its incredibly promising future, long read sequencing does not come without its own set of challenges. In this thesis, I explore several different applications of long read DNA sequencing and improve upon current methodologies in this new field. First, we present a hardware-accelerated filter that directly analyzes nanopore sequencer output in real time to filter non-viral reads, enabling cheaper detection of pathogenic viruses. Next, we introduce a novel read alignment algorithm that enables more consistent alignment of long reads in highly repetitive areas of the genome, and demonstrate that this improves recall for tandem repeat variant calling. Then, we analyze the design space for complex variant representation and present a new variant calling benchmarking tool that accurately and stably measures performance regardless of the representation of reported variants. Last, we extend this benchmarking tool to jointly evaluate small and structural variants, and demonstrate that doing so results in improved measured performance and enables more accurate phasing analyses.
일반주제명  
Computer science
일반주제명  
Bioinformatics
일반주제명  
Molecular biology
일반주제명  
Genetics
키워드  
Long read sequencing
키워드  
Nanopore sequencing
키워드  
Whole genome sequencing
키워드  
Alignment
키워드  
Variant calling
키워드  
Benchmarking
기타저자  
University of Michigan Computer Science & Engineering
기본자료저록  
Dissertations Abstracts International. 86-03B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017164355
■00520250211152952
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798384042013
■035    ▼a(MiAaPQ)AAI31631054
■035    ▼a(MiAaPQ)umichrackham005838
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aDunn,  Timothy  J.
■24510▼aImproving  Select  Applications  of  Long-Read  DNA  Sequencing
■260    ▼a[Sl]▼bUniversity  of  Michigan▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a184  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-03,  Section:  B.
■500    ▼aAdvisor:  Narayanasamy,  Satish.
■5021  ▼aThesis  (Ph.D.)--University  of  Michigan,  2024.
■520    ▼aThe  cost  to  sequence  a  human  genome  has  dropped  from  an  estimated  $300  million  to  well  under  $1,000  over  the  past  two  decades.  In  fact,  several  companies  -  which  can  amortize  sequencing  cost  by  running  hundreds  of  samples  in  parallel  -  have  recently  claimed  to  have  reached  the  $100  human  genome.  Following  such  dramatic  cost  improvements,  whole  genome  sequencing  is  just  starting  to  regularly  be  used  for  cancer  profiling,  rare  genetic  disease  detection,  agricultural  breeding,  pathogen  detection,  microbiome  bacterial  abundance  estimation,  evolutionary  biology  research,  personalized  medicine  development,  and  much  more.  As  sequencing  costs  continue  to  drop,  as  accuracy  improves,  and  as  new  applications  are  discovered,  DNA  sequencing  will  become  increasingly  ubiquitous.  At  the  same  time  that  sequencing  cost  is  rapidly  dropping,  new  sequencing  technologies  have  emerged  that  offer  greater  capabilities  than  ever  before.  In  particular,  nanopore-based  long-read  sequencing  has  no  theoretical  limit  on  the  length  of  a  contiguous  DNA  sequence,  or  "read",  that  can  be  measured.  In  comparison  to  short-read  sequencing  technologies  that  have  dominated  the  sequencing  market  thus  far  (with  maximum  read  lengths  of  100  base  pairs),  nanopore  devices  have  sequenced  entire  bacterial  chromosomes  in  a  single  strand.  The  current  nanopore  read  length  record  stands  at  over  four  million  bases.  Longer  read  lengths  result  in  fewer  problems  during  read  mapping  and  genome  assembly,  allowing  insight  into  complex  regions  of  the  genome  and  types  of  genetic  variation  that  have  been  historically  under-studied.  Although  nanopore  devices  were  originally  limited  by  their  approximately  80%  per-base  accuracy  when  first  publicly  released  in  2015,  this  accuracy  has  increased  to  over  99%  in  recent  years  with  the  adoption  of  deep  learning  basecallers.  Nanopore-based  sequencing  devices  are  also  the  first  to  come  in  a  portable  handheld  form  factor  and  offer  real-time  analysis  of  raw  data  as  it  is  being  recorded,  further  expanding  potential  use  cases.  Despite  its  incredibly  promising  future,  long  read  sequencing  does  not  come  without  its  own  set  of  challenges.  In  this  thesis,  I  explore  several  different  applications  of  long  read  DNA  sequencing  and  improve  upon  current  methodologies  in  this  new  field.  First,  we  present  a  hardware-accelerated  filter  that  directly  analyzes  nanopore  sequencer  output  in  real  time  to  filter  non-viral  reads,  enabling  cheaper  detection  of  pathogenic  viruses.  Next,  we  introduce  a  novel  read  alignment  algorithm  that  enables  more  consistent  alignment  of  long  reads  in  highly  repetitive  areas  of  the  genome,  and  demonstrate  that  this  improves  recall  for  tandem  repeat  variant  calling.  Then,  we  analyze  the  design  space  for  complex  variant  representation  and  present  a  new  variant  calling  benchmarking  tool  that  accurately  and  stably  measures  performance  regardless  of  the  representation  of  reported  variants.  Last,  we  extend  this  benchmarking  tool  to  jointly  evaluate  small  and  structural  variants,  and  demonstrate  that  doing  so  results  in  improved  measured  performance  and  enables  more  accurate  phasing  analyses.
■590    ▼aSchool  code:  0127.
■650  4▼aComputer  science
■650  4▼aBioinformatics
■650  4▼aMolecular  biology
■650  4▼aGenetics
■653    ▼aLong  read  sequencing
■653    ▼aNanopore  sequencing
■653    ▼aWhole  genome  sequencing
■653    ▼aAlignment
■653    ▼aVariant  calling
■653    ▼aBenchmarking
■690    ▼a0984
■690    ▼a0715
■690    ▼a0369
■690    ▼a0307
■71020▼aUniversity  of  Michigan▼bComputer  Science  &  Engineering.
■7730  ▼tDissertations  Abstracts  International▼g86-03B.
■790    ▼a0127
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164355▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF10985 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.