서브메뉴
검색
Machine Learning Approaches For Predicting Biological Function and Molecular Design
Machine Learning Approaches For Predicting Biological Function and Molecular Design
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103138
- ISBN
- 9798311957991
- DDC
- 541.39
- 서명/저자
- Machine Learning Approaches For Predicting Biological Function and Molecular Design
- 발행사항
- [Sl] : Stanford University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 292 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-12, Section: B.
- 주기사항
- Advisor: Bassik, Michael;Kundaje, Anshul.
- 학위논문주기
- Thesis (Ph.D.)--Stanford University, 2025.
- 초록/해제
- 요약Drug discovery represents one of the most intellectually demanding and resource-intensive challenges in modern science. On average, it costs approximately $1.1 billion to develop a single drug, and this process can span over a decade. These staggering figures reflect the complexity of identifying, validating, and optimizing compounds that can e↵ectively target specific biological mechanisms without causing unacceptable side e↵ects. Compounding these challenges, the success rate of drugs making it through clinical trials is less than 10%, underscoring the ineciency and unpredictability of the traditional drug development pipeline.Despite significant advancements in computational tools and experimental techniques, the process of discovering small molecules and biologics remains constrained by the scale of chemical and biological complexity involved. As a result, there is an urgent need to develop innovative computational frameworks that can improve the accuracy, scalability, and eciency of key steps in drug discovery. This doctoral research addresses these challenges through three major areas: the prediction and design of transcriptional repressors, the application of ultra-large virtual screening (ULVS) for small molecule discovery, and the creation of robust benchmarks for generative molecular design models. Together, these e↵orts aim to redefine the boundaries of computational drug discovery by integrating machine learning with experimental validation.1.1 Predicting and Designing Transcriptional RepressorsTranscriptional repressors play a pivotal role in regulating gene expression by silencing specific genes in response to biological signals. They are critical in maintaining cellular homeostasis and have profound implications in diseases such as cancer, autoimmune disorders, and neurodegeneration. However, the design of novel transcriptional repressors and the understanding of how specific mutations a↵ect their function remain significant challenges.This research focuses on harnessing deep mutational scanning (DMS) data to predict and design transcriptional repressors with enhanced functionality. DMS provides an exhaustive mapping of mutational e↵ects across protein domains, o↵ering a treasure trove of information for data-driven approaches. To capitalize on this, I developed the Transcriptional E↵ector Network (TENet)[424], a state-of-the-art deep learning model that integrates sequence-based embeddings, amino acid descriptors, and structural features. TENet achieves high predictive accuracy and generalizability across diverse protein domains, enabling not only the prediction of repressor activity but also the rational design of novel repressor variants. This work bridges computational modeling and experimental validation, providing insights into the design principles of transcriptional e↵ectors.1.2 Ultra-Large Virtual Screening for Small Molecule DiscoverySmall molecules account for the majority of FDA-approved drugs and remain a cornerstone of therapeutic innovation. However, the chemical space of potential small molecules is vast, estimated to be between 1060and 10100compounds, far exceeding the number of molecules that can be synthesized or experimentally screened. Virtual screening, a computational technique to identify promising candidates from large libraries of compounds, o↵ers a solution to this bottleneck. However, traditional virtual screening approaches are limited in scale and often lack the computational eciency to screen billions of compounds.To address this, I contributed to the development of VirtualFlow 2.0[126], a next-generation platform for ultra-large virtual screening (ULVS). VirtualFlow 2.0 introduces adaptive screening algorithms that prioritize high-confidence compounds during docking simulations, significantly improving eciency without compromising accuracy. Additionally, the platform is designed to be infrastructure-agnostic, enabling seamless deployment on academic clusters, cloud computing services, and other high-performance computing environments. Through VirtualFlow 2.0, I applied ULVS to identify novel small molecules targeting challenging proteins, including PARP1, a critical oncogene implicated in several cancers. These e↵orts demonstrate the power of ULVS to democratize access to large-scale computational drug discovery and accelerate the identification of promising therapeutic candidates.1.3 Benchmarking Generative Models for Molecular DesignThe rise of generative models in molecular design has opened new possibilities for exploring chemical space and optimizing compounds for specific properties. Generative models, powered by advances in machine learning, can propose novel molecules with desired characteristics, o↵ering a complementary approach to traditional drug discovery methods. However, the evaluation of generative models has often relied on oversimplified benchmarks that do not reflect the complexities of real-world drug discovery tasks.To address this gap, I co-developed Tartarus, [301] a comprehensive benchmarking platform for generative molecular models. Tartarus introduces datasets and metrics that better capture the practical challenges faced in molecular design, such as synthetic accessibility, property optimization, and multi-objective trade-o↵s. One of the key contributions of this work is the application of Tartarus to hybrid quantum-classical generative models. By leveraging quantum computing's ability to explore high-dimensional parameter spaces, I demonstrated the potential of these models to generate promising compounds, including novel KRAS inhibitors. [422] This work underscores the importance of realistic and rigorous benchmarks in driving progress in molecular generative modeling.1.4 Broader Implications and Future DirectionsA key theme across this research is the integration of computational predictions with experimental validation. Computational approaches, while powerful, are only as impactful as their ability to guide meaningful experiments. In my work on transcriptional repressors, predictions from TENet were experimentally validated through deep mutational scanning and functional assays. Similarly, compounds identified through VirtualFlow 2.0 underwent synthesis and biological testing, providing critical feedback for refining computational models.This iterative cycle of prediction, validation, and refinement exemplifies a modern approach to drug discovery, where computation and experimentation work hand in hand. By bridging these domains, this research has contributed to the development of more reliable and interpretable models that can accelerate the translation of computational insights into tangible therapeutic advancements.
- 일반주제명
- Chemical reactions
- 일반주제명
- Neural networks
- 일반주제명
- Benchmarks
- 일반주제명
- Bioinformatics
- 기타저자
- Stanford University.
- 기본자료저록
- Dissertations Abstracts International. 86-12B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017357142
■00520260202103138
■006m o d
■007cr#unu||||||||
■020 ▼a9798311957991
■035 ▼a(MiAaPQ)AAI31974665
■035 ▼a(MiAaPQ)Stanfordth395fw7851
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a541.39
■1001 ▼aNigam, Akshatkumar.
■24510▼aMachine Learning Approaches For Predicting Biological Function and Molecular Design
■260 ▼a[Sl]▼bStanford University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a292 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-12, Section: B.
■500 ▼aAdvisor: Bassik, Michael;Kundaje, Anshul.
■5021 ▼aThesis (Ph.D.)--Stanford University, 2025.
■520 ▼aDrug discovery represents one of the most intellectually demanding and resource-intensive challenges in modern science. On average, it costs approximately $1.1 billion to develop a single drug, and this process can span over a decade. These staggering figures reflect the complexity of identifying, validating, and optimizing compounds that can e↵ectively target specific biological mechanisms without causing unacceptable side e↵ects. Compounding these challenges, the success rate of drugs making it through clinical trials is less than 10%, underscoring the ineciency and unpredictability of the traditional drug development pipeline.Despite significant advancements in computational tools and experimental techniques, the process of discovering small molecules and biologics remains constrained by the scale of chemical and biological complexity involved. As a result, there is an urgent need to develop innovative computational frameworks that can improve the accuracy, scalability, and eciency of key steps in drug discovery. This doctoral research addresses these challenges through three major areas: the prediction and design of transcriptional repressors, the application of ultra-large virtual screening (ULVS) for small molecule discovery, and the creation of robust benchmarks for generative molecular design models. Together, these e↵orts aim to redefine the boundaries of computational drug discovery by integrating machine learning with experimental validation.1.1 Predicting and Designing Transcriptional RepressorsTranscriptional repressors play a pivotal role in regulating gene expression by silencing specific genes in response to biological signals. They are critical in maintaining cellular homeostasis and have profound implications in diseases such as cancer, autoimmune disorders, and neurodegeneration. However, the design of novel transcriptional repressors and the understanding of how specific mutations a↵ect their function remain significant challenges.This research focuses on harnessing deep mutational scanning (DMS) data to predict and design transcriptional repressors with enhanced functionality. DMS provides an exhaustive mapping of mutational e↵ects across protein domains, o↵ering a treasure trove of information for data-driven approaches. To capitalize on this, I developed the Transcriptional E↵ector Network (TENet)[424], a state-of-the-art deep learning model that integrates sequence-based embeddings, amino acid descriptors, and structural features. TENet achieves high predictive accuracy and generalizability across diverse protein domains, enabling not only the prediction of repressor activity but also the rational design of novel repressor variants. This work bridges computational modeling and experimental validation, providing insights into the design principles of transcriptional e↵ectors.1.2 Ultra-Large Virtual Screening for Small Molecule DiscoverySmall molecules account for the majority of FDA-approved drugs and remain a cornerstone of therapeutic innovation. However, the chemical space of potential small molecules is vast, estimated to be between 1060and 10100compounds, far exceeding the number of molecules that can be synthesized or experimentally screened. Virtual screening, a computational technique to identify promising candidates from large libraries of compounds, o↵ers a solution to this bottleneck. However, traditional virtual screening approaches are limited in scale and often lack the computational eciency to screen billions of compounds.To address this, I contributed to the development of VirtualFlow 2.0[126], a next-generation platform for ultra-large virtual screening (ULVS). VirtualFlow 2.0 introduces adaptive screening algorithms that prioritize high-confidence compounds during docking simulations, significantly improving eciency without compromising accuracy. Additionally, the platform is designed to be infrastructure-agnostic, enabling seamless deployment on academic clusters, cloud computing services, and other high-performance computing environments. Through VirtualFlow 2.0, I applied ULVS to identify novel small molecules targeting challenging proteins, including PARP1, a critical oncogene implicated in several cancers. These e↵orts demonstrate the power of ULVS to democratize access to large-scale computational drug discovery and accelerate the identification of promising therapeutic candidates.1.3 Benchmarking Generative Models for Molecular DesignThe rise of generative models in molecular design has opened new possibilities for exploring chemical space and optimizing compounds for specific properties. Generative models, powered by advances in machine learning, can propose novel molecules with desired characteristics, o↵ering a complementary approach to traditional drug discovery methods. However, the evaluation of generative models has often relied on oversimplified benchmarks that do not reflect the complexities of real-world drug discovery tasks.To address this gap, I co-developed Tartarus, [301] a comprehensive benchmarking platform for generative molecular models. Tartarus introduces datasets and metrics that better capture the practical challenges faced in molecular design, such as synthetic accessibility, property optimization, and multi-objective trade-o↵s. One of the key contributions of this work is the application of Tartarus to hybrid quantum-classical generative models. By leveraging quantum computing's ability to explore high-dimensional parameter spaces, I demonstrated the potential of these models to generate promising compounds, including novel KRAS inhibitors. [422] This work underscores the importance of realistic and rigorous benchmarks in driving progress in molecular generative modeling.1.4 Broader Implications and Future DirectionsA key theme across this research is the integration of computational predictions with experimental validation. Computational approaches, while powerful, are only as impactful as their ability to guide meaningful experiments. In my work on transcriptional repressors, predictions from TENet were experimentally validated through deep mutational scanning and functional assays. Similarly, compounds identified through VirtualFlow 2.0 underwent synthesis and biological testing, providing critical feedback for refining computational models.This iterative cycle of prediction, validation, and refinement exemplifies a modern approach to drug discovery, where computation and experimentation work hand in hand. By bridging these domains, this research has contributed to the development of more reliable and interpretable models that can accelerate the translation of computational insights into tangible therapeutic advancements.
■590 ▼aSchool code: 0212.
■650 4▼aChemical reactions
■650 4▼aNeural networks
■650 4▼aBenchmarks
■650 4▼aBioinformatics
■690 ▼a0715
■71020▼aStanford University.
■7730 ▼tDissertations Abstracts International▼g86-12B.
■790 ▼a0212
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357142▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


