서브메뉴
검색
Improving Select Applications of Long-Read DNA Sequencing
Improving Select Applications of Long-Read DNA Sequencing
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211152952
- ISBN
- 9798384042013
- DDC
- 004
- 저자명
- Dunn, Timothy J.
- 서명/저자
- Improving Select Applications of Long-Read DNA Sequencing
- 발행사항
- [Sl] : University of Michigan, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 184 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-03, Section: B.
- 주기사항
- Advisor: Narayanasamy, Satish.
- 학위논문주기
- Thesis (Ph.D.)--University of Michigan, 2024.
- 초록/해제
- 요약The cost to sequence a human genome has dropped from an estimated $300 million to well under $1,000 over the past two decades. In fact, several companies - which can amortize sequencing cost by running hundreds of samples in parallel - have recently claimed to have reached the $100 human genome. Following such dramatic cost improvements, whole genome sequencing is just starting to regularly be used for cancer profiling, rare genetic disease detection, agricultural breeding, pathogen detection, microbiome bacterial abundance estimation, evolutionary biology research, personalized medicine development, and much more. As sequencing costs continue to drop, as accuracy improves, and as new applications are discovered, DNA sequencing will become increasingly ubiquitous. At the same time that sequencing cost is rapidly dropping, new sequencing technologies have emerged that offer greater capabilities than ever before. In particular, nanopore-based long-read sequencing has no theoretical limit on the length of a contiguous DNA sequence, or "read", that can be measured. In comparison to short-read sequencing technologies that have dominated the sequencing market thus far (with maximum read lengths of 100 base pairs), nanopore devices have sequenced entire bacterial chromosomes in a single strand. The current nanopore read length record stands at over four million bases. Longer read lengths result in fewer problems during read mapping and genome assembly, allowing insight into complex regions of the genome and types of genetic variation that have been historically under-studied. Although nanopore devices were originally limited by their approximately 80% per-base accuracy when first publicly released in 2015, this accuracy has increased to over 99% in recent years with the adoption of deep learning basecallers. Nanopore-based sequencing devices are also the first to come in a portable handheld form factor and offer real-time analysis of raw data as it is being recorded, further expanding potential use cases. Despite its incredibly promising future, long read sequencing does not come without its own set of challenges. In this thesis, I explore several different applications of long read DNA sequencing and improve upon current methodologies in this new field. First, we present a hardware-accelerated filter that directly analyzes nanopore sequencer output in real time to filter non-viral reads, enabling cheaper detection of pathogenic viruses. Next, we introduce a novel read alignment algorithm that enables more consistent alignment of long reads in highly repetitive areas of the genome, and demonstrate that this improves recall for tandem repeat variant calling. Then, we analyze the design space for complex variant representation and present a new variant calling benchmarking tool that accurately and stably measures performance regardless of the representation of reported variants. Last, we extend this benchmarking tool to jointly evaluate small and structural variants, and demonstrate that doing so results in improved measured performance and enables more accurate phasing analyses.
- 일반주제명
- Computer science
- 일반주제명
- Bioinformatics
- 일반주제명
- Molecular biology
- 일반주제명
- Genetics
- 키워드
- Alignment
- 키워드
- Variant calling
- 키워드
- Benchmarking
- 기타저자
- University of Michigan Computer Science & Engineering
- 기본자료저록
- Dissertations Abstracts International. 86-03B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017164355
■00520250211152952
■006m o d
■007cr#unu||||||||
■020 ▼a9798384042013
■035 ▼a(MiAaPQ)AAI31631054
■035 ▼a(MiAaPQ)umichrackham005838
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aDunn, Timothy J.
■24510▼aImproving Select Applications of Long-Read DNA Sequencing
■260 ▼a[Sl]▼bUniversity of Michigan▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a184 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-03, Section: B.
■500 ▼aAdvisor: Narayanasamy, Satish.
■5021 ▼aThesis (Ph.D.)--University of Michigan, 2024.
■520 ▼aThe cost to sequence a human genome has dropped from an estimated $300 million to well under $1,000 over the past two decades. In fact, several companies - which can amortize sequencing cost by running hundreds of samples in parallel - have recently claimed to have reached the $100 human genome. Following such dramatic cost improvements, whole genome sequencing is just starting to regularly be used for cancer profiling, rare genetic disease detection, agricultural breeding, pathogen detection, microbiome bacterial abundance estimation, evolutionary biology research, personalized medicine development, and much more. As sequencing costs continue to drop, as accuracy improves, and as new applications are discovered, DNA sequencing will become increasingly ubiquitous. At the same time that sequencing cost is rapidly dropping, new sequencing technologies have emerged that offer greater capabilities than ever before. In particular, nanopore-based long-read sequencing has no theoretical limit on the length of a contiguous DNA sequence, or "read", that can be measured. In comparison to short-read sequencing technologies that have dominated the sequencing market thus far (with maximum read lengths of 100 base pairs), nanopore devices have sequenced entire bacterial chromosomes in a single strand. The current nanopore read length record stands at over four million bases. Longer read lengths result in fewer problems during read mapping and genome assembly, allowing insight into complex regions of the genome and types of genetic variation that have been historically under-studied. Although nanopore devices were originally limited by their approximately 80% per-base accuracy when first publicly released in 2015, this accuracy has increased to over 99% in recent years with the adoption of deep learning basecallers. Nanopore-based sequencing devices are also the first to come in a portable handheld form factor and offer real-time analysis of raw data as it is being recorded, further expanding potential use cases. Despite its incredibly promising future, long read sequencing does not come without its own set of challenges. In this thesis, I explore several different applications of long read DNA sequencing and improve upon current methodologies in this new field. First, we present a hardware-accelerated filter that directly analyzes nanopore sequencer output in real time to filter non-viral reads, enabling cheaper detection of pathogenic viruses. Next, we introduce a novel read alignment algorithm that enables more consistent alignment of long reads in highly repetitive areas of the genome, and demonstrate that this improves recall for tandem repeat variant calling. Then, we analyze the design space for complex variant representation and present a new variant calling benchmarking tool that accurately and stably measures performance regardless of the representation of reported variants. Last, we extend this benchmarking tool to jointly evaluate small and structural variants, and demonstrate that doing so results in improved measured performance and enables more accurate phasing analyses.
■590 ▼aSchool code: 0127.
■650 4▼aComputer science
■650 4▼aBioinformatics
■650 4▼aMolecular biology
■650 4▼aGenetics
■653 ▼aLong read sequencing
■653 ▼aNanopore sequencing
■653 ▼aWhole genome sequencing
■653 ▼aAlignment
■653 ▼aVariant calling
■653 ▼aBenchmarking
■690 ▼a0984
■690 ▼a0715
■690 ▼a0369
■690 ▼a0307
■71020▼aUniversity of Michigan▼bComputer Science & Engineering.
■7730 ▼tDissertations Abstracts International▼g86-03B.
■790 ▼a0127
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164355▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


