서브메뉴
검색
From Arabidopsis to Zea: Learning Conserved Cis Mechanisms of Gene Regulation
From Arabidopsis to Zea: Learning Conserved Cis Mechanisms of Gene Regulation
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211152716
- ISBN
- 9798384053590
- DDC
- 580
- 서명/저자
- From Arabidopsis to Zea: Learning Conserved Cis Mechanisms of Gene Regulation
- 발행사항
- [Sl] : Cornell University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 84 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-03, Section: B.
- 주기사항
- Advisor: Buckler, Edward.
- 학위논문주기
- Thesis (Ph.D.)--Cornell University, 2024.
- 초록/해제
- 요약cis-regulatory elements (CREs) are critical functional components of the genome, controlling both the timing and magnitude of gene expression. Relative to protein-coding regions, CREs have proven more difficult and expensive to catalogue even within humans, with CRE knowledge in other higher organisms trailing far behind. To reduce the cost of locating CREs, many deep learning sequence-based models have been developed within species to predict CREs directly from DNA sequence. However, the vast majority of these models are validated within species or against species in the training set and never tested on a completely held-out set of species, questioning their generalizability. This dissertation explores three methods to locate CREs in held-out species at different phylogenetic scopes, from angiosperms to the Andropogoneae. The first method uses a recurrent convolutional neural network to classify 600 base pair sequence windows as accessible or hypomethylated in leaf tissue, two epigenetic signals associated with CREs. Models trained across multiple species and tested on a held-out species perform competitively or superior to models trained and tested solely within species. These multispecies models demonstrate the feasibility of predicting epigenetic signals of CREs in understudied species. The second method compares four published genomic deep learning model architectures on their ability to predict RNA abundance from 1,000 base pairs of promoter and UTR sequence. Models trained across the Andropogoneae and tested within maize showed moderate performance across all genes but poor performance within maize orthogroups. The dataset used to fairly compare all architectures has been publicly released as a community resource to consistently benchmark future expression model architectures. The final method uses phylogenetic footprinting with hundreds of Andropogoneae genomes to filter motif matches in maize to likely functional transcription factor binding sites. Aided by the high alignment depth, motifs within some transcription factor families show strong clustering into novel subfamily motifs that can be associated with changes in tissue-specific gene expression. These novel subfamilies are promising candidates for development and stress-specific transcription factor family members. Together, these three methods demonstrate the utility in leveraging data from many related species to identify CREs or functional loci within CREs, which can be useful targets for genome editing for crop improvement.
- 일반주제명
- Plant sciences
- 일반주제명
- Bioinformatics
- 일반주제명
- Genetics
- 키워드
- Andropogoneae
- 키워드
- DNA sequence
- 키워드
- Epigenetic
- 키워드
- Genome editing
- 키워드
- Crop improvement
- 기타저자
- Cornell University Plant Breeding
- 기본자료저록
- Dissertations Abstracts International. 86-03B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017163504
■00520250211152716
■006m o d
■007cr#unu||||||||
■020 ▼a9798384053590
■035 ▼a(MiAaPQ)AAI31489147
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a580
■1001 ▼aWrightsman, Travis.▼0(orcid)0000-0002-0904-6473
■24510▼aFrom Arabidopsis to Zea: Learning Conserved Cis Mechanisms of Gene Regulation
■260 ▼a[Sl]▼bCornell University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a84 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-03, Section: B.
■500 ▼aAdvisor: Buckler, Edward.
■5021 ▼aThesis (Ph.D.)--Cornell University, 2024.
■520 ▼acis-regulatory elements (CREs) are critical functional components of the genome, controlling both the timing and magnitude of gene expression. Relative to protein-coding regions, CREs have proven more difficult and expensive to catalogue even within humans, with CRE knowledge in other higher organisms trailing far behind. To reduce the cost of locating CREs, many deep learning sequence-based models have been developed within species to predict CREs directly from DNA sequence. However, the vast majority of these models are validated within species or against species in the training set and never tested on a completely held-out set of species, questioning their generalizability. This dissertation explores three methods to locate CREs in held-out species at different phylogenetic scopes, from angiosperms to the Andropogoneae. The first method uses a recurrent convolutional neural network to classify 600 base pair sequence windows as accessible or hypomethylated in leaf tissue, two epigenetic signals associated with CREs. Models trained across multiple species and tested on a held-out species perform competitively or superior to models trained and tested solely within species. These multispecies models demonstrate the feasibility of predicting epigenetic signals of CREs in understudied species. The second method compares four published genomic deep learning model architectures on their ability to predict RNA abundance from 1,000 base pairs of promoter and UTR sequence. Models trained across the Andropogoneae and tested within maize showed moderate performance across all genes but poor performance within maize orthogroups. The dataset used to fairly compare all architectures has been publicly released as a community resource to consistently benchmark future expression model architectures. The final method uses phylogenetic footprinting with hundreds of Andropogoneae genomes to filter motif matches in maize to likely functional transcription factor binding sites. Aided by the high alignment depth, motifs within some transcription factor families show strong clustering into novel subfamily motifs that can be associated with changes in tissue-specific gene expression. These novel subfamilies are promising candidates for development and stress-specific transcription factor family members. Together, these three methods demonstrate the utility in leveraging data from many related species to identify CREs or functional loci within CREs, which can be useful targets for genome editing for crop improvement.
■590 ▼aSchool code: 0058.
■650 4▼aPlant sciences
■650 4▼aBioinformatics
■650 4▼aGenetics
■653 ▼aAndropogoneae
■653 ▼aCis-regulatory elements
■653 ▼aDNA sequence
■653 ▼aEpigenetic
■653 ▼aGenome editing
■653 ▼aCrop improvement
■690 ▼a0479
■690 ▼a0715
■690 ▼a0369
■71020▼aCornell University▼bPlant Breeding.
■7730 ▼tDissertations Abstracts International▼g86-03B.
■790 ▼a0058
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17163504▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


