서브메뉴
검색
Foundations for Genome-Scale Artificial Intelligence
Foundations for Genome-Scale Artificial Intelligence
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103559
- ISBN
- 9798280712843
- DDC
- 574
- 저자명
- Ektefaie, Yasha.
- 서명/저자
- Foundations for Genome-Scale Artificial Intelligence
- 발행사항
- [Sl] : Harvard University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 233 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-12, Section: B.
- 주기사항
- Advisor: Farhat, Maha;Zitnik, Marinka.
- 학위논문주기
- Thesis (Ph.D.)--Harvard University, 2025.
- 초록/해제
- 요약Modeling whole genomes-the complete sequence of base pairs, including both coding and non-coding regions organized within a three-dimensional architecture-represents a grand challenge in bioinformatics. Achieving this goal could transform our ability to understand complex polygenic diseases and develop treatments for both rare and common conditions. However, whole-genome modeling presents two fundamental challenges: the vast size of genomic data and the intricate complexity of genetic interactions. Artificial intelligence (AI) offers a powerful approach to processing large datasets and extracting meaningful patterns. However, for AI to effectively model whole genomes, it must be (1) generalizable, capable of predicting the effects of novel, unseen mutations; (2) capable of reasoning across sequences and biological scales, as genetic function emerges from the multi-level interactions of sequences; and (3) multimodal, integrating information from diverse sources, including scientific literature. While this dissertation does not yet achieve full-scale genome modeling, it introduces four novel AI methodologies-SPECTRA, Phyla, RLDIF, and Fleming-each addressing essential components necessary to enable this vision. Together, these contributions establish a conceptual and methodological framework that lays the groundwork for the future development of genome-scale artificial intelligence. SPECTRA is a framework for evaluating model generalizability beyond conventional dataset splits. By systematically varying train-test similarity, SPECTRA reveals that existing biological foundation models fail to generalize to sequences dissimilar from their training data. In response, Phyla is designed to explicitly learn how to compare sequences by leveraging evolutionary relationships. Trained on protein phylogenies, Phyla not only excels at sequence comparison but also reconstructs phylogenetic trees with high accuracy, revealing both known and novel evolutionary insights. Additionally, this dissertation presents RLDIF, a categorical conditional diffusion model for protein inverse folding, which leverages reinforcement learning to improve multi-scale modeling. By optimizing sequence design with respect to structural recovery, RLDIF achieves state-of-the-art performance in generating diverse sequences that accurately fold into a specified target structure. Beyond sequence information, AI models must incorporate knowledge from external sources to avoid rediscovering established principles. To address this, Fleming is developed as an AI agent for antibiotic design in tuberculosis, integrating scientific literature with machine learning tools to generate novel antibiotic candidates. The same multimodal approach can be extended to genome analysis, enabling AI to reason over diverse data modalities. Together, these innovations establish a foundation for genome-scale AI, providing key tools for understanding and reasoning over whole genomes.
- 일반주제명
- Bioinformatics
- 일반주제명
- Genetics
- 키워드
- Machine learning
- 키워드
- Genomes
- 기타저자
- Harvard University Biomedical Informatics
- 기본자료저록
- Dissertations Abstracts International. 86-12B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017357778
■00520260202103559
■006m o d
■007cr#unu||||||||
■020 ▼a9798280712843
■035 ▼a(MiAaPQ)AAI32042299
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a574
■1001 ▼aEktefaie, Yasha.▼0(orcid)0000-0003-2759-4470
■24510▼aFoundations for Genome-Scale Artificial Intelligence
■260 ▼a[Sl]▼bHarvard University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a233 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-12, Section: B.
■500 ▼aAdvisor: Farhat, Maha;Zitnik, Marinka.
■5021 ▼aThesis (Ph.D.)--Harvard University, 2025.
■520 ▼aModeling whole genomes-the complete sequence of base pairs, including both coding and non-coding regions organized within a three-dimensional architecture-represents a grand challenge in bioinformatics. Achieving this goal could transform our ability to understand complex polygenic diseases and develop treatments for both rare and common conditions. However, whole-genome modeling presents two fundamental challenges: the vast size of genomic data and the intricate complexity of genetic interactions. Artificial intelligence (AI) offers a powerful approach to processing large datasets and extracting meaningful patterns. However, for AI to effectively model whole genomes, it must be (1) generalizable, capable of predicting the effects of novel, unseen mutations; (2) capable of reasoning across sequences and biological scales, as genetic function emerges from the multi-level interactions of sequences; and (3) multimodal, integrating information from diverse sources, including scientific literature. While this dissertation does not yet achieve full-scale genome modeling, it introduces four novel AI methodologies-SPECTRA, Phyla, RLDIF, and Fleming-each addressing essential components necessary to enable this vision. Together, these contributions establish a conceptual and methodological framework that lays the groundwork for the future development of genome-scale artificial intelligence. SPECTRA is a framework for evaluating model generalizability beyond conventional dataset splits. By systematically varying train-test similarity, SPECTRA reveals that existing biological foundation models fail to generalize to sequences dissimilar from their training data. In response, Phyla is designed to explicitly learn how to compare sequences by leveraging evolutionary relationships. Trained on protein phylogenies, Phyla not only excels at sequence comparison but also reconstructs phylogenetic trees with high accuracy, revealing both known and novel evolutionary insights. Additionally, this dissertation presents RLDIF, a categorical conditional diffusion model for protein inverse folding, which leverages reinforcement learning to improve multi-scale modeling. By optimizing sequence design with respect to structural recovery, RLDIF achieves state-of-the-art performance in generating diverse sequences that accurately fold into a specified target structure. Beyond sequence information, AI models must incorporate knowledge from external sources to avoid rediscovering established principles. To address this, Fleming is developed as an AI agent for antibiotic design in tuberculosis, integrating scientific literature with machine learning tools to generate novel antibiotic candidates. The same multimodal approach can be extended to genome analysis, enabling AI to reason over diverse data modalities. Together, these innovations establish a foundation for genome-scale AI, providing key tools for understanding and reasoning over whole genomes.
■590 ▼aSchool code: 0084.
■650 4▼aBioinformatics
■650 4▼aGenetics
■653 ▼aMachine learning
■653 ▼aGenomes
■653 ▼aEvolutionary relationships
■653 ▼aPhylogenetic trees
■690 ▼a0715
■690 ▼a0800
■690 ▼a0369
■71020▼aHarvard University▼bBiomedical Informatics.
■7730 ▼tDissertations Abstracts International▼g86-12B.
■790 ▼a0084
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357778▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


