서브메뉴
검색
Autonomous Experiment Design in Chemistry With Machine Learning
Autonomous Experiment Design in Chemistry With Machine Learning
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202104821
- ISBN
- 9798290938639
- DDC
- 660
- 저자명
- Boiko, Daniil A.
- 서명/저자
- Autonomous Experiment Design in Chemistry With Machine Learning
- 발행사항
- [Sl] : Carnegie Mellon University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 271 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-02, Section: B.
- 주기사항
- Advisor: Gomes, Gabe.
- 학위논문주기
- Thesis (Ph.D.)--Carnegie Mellon University, 2025.
- 초록/해제
- 요약Recent advances in data‐driven science promise to transform chemical discovery, yet most machine-learning tools remain siloed, dataset-limited, or disconnected from laboratory workflows. This thesis proposes an end-to-end framework that closes that gap by (i) inventing richer molecular representations, (ii) developing new methods for enzyme-substrate activity prediction task, (iii) scaling experimental data generation, and (iv) embedding those assets in an autonomous "co-scientist" platform.First, we introduce stereoelectronics-infused molecular graphs (SIMGs)-hypergraphs that encode lone pairs, bond orbitals and orbital interactions extracted from Natural Bond Orbital (NBO) analysis. A two-stage graph-neural pipeline approximates SIMGs, delivering high accuracy while remaining tractable for large molecules.Second, the work tackles biocatalysis, where enzyme-substrate activity is notoriously hard to predict. Leveraging ranking models, we build a recommender system that prioritizes enzymes for novel substrates.Third, we address data scarcity by constructing the largest experimental reaction dataset to date. High-resolution flow-injection MS, robotically prepared imine libraries, and reaction-based multiplexing yield 50,000 sample-compound pairs, capturing full kinetic and compositional profiles indispensable for next-generation ML models.Finally, these components integrate into Coscientist, an LLM-orchestrated agent that plans, executes and interprets experiments across heterogeneous lab hardware. Validation spans autonomous optimization of Pd-catalyzed couplings and multi-tool instrument control, illustrating how the platform can compress the design-make-test-analyze cycle.Collectively, the thesis advances molecular representation, dataset scale, and laboratory autonomy, laying a foundation for ML-driven experimentation in chemistry.
- 일반주제명
- Chemical engineering
- 일반주제명
- Materials science
- 일반주제명
- Molecular chemistry
- 키워드
- Biocatalysis
- 기타저자
- Carnegie Mellon University Chemical Engineering
- 기본자료저록
- Dissertations Abstracts International. 87-02B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017359009
■00520260202104821
■006m o d
■007cr#unu||||||||
■020 ▼a9798290938639
■035 ▼a(MiAaPQ)AAI32169202
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a660
■1001 ▼aBoiko, Daniil A.▼0(orcid)0000-0003-4140-4645
■24510▼aAutonomous Experiment Design in Chemistry With Machine Learning
■260 ▼a[Sl]▼bCarnegie Mellon University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a271 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-02, Section: B.
■500 ▼aAdvisor: Gomes, Gabe.
■5021 ▼aThesis (Ph.D.)--Carnegie Mellon University, 2025.
■520 ▼aRecent advances in data‐driven science promise to transform chemical discovery, yet most machine-learning tools remain siloed, dataset-limited, or disconnected from laboratory workflows. This thesis proposes an end-to-end framework that closes that gap by (i) inventing richer molecular representations, (ii) developing new methods for enzyme-substrate activity prediction task, (iii) scaling experimental data generation, and (iv) embedding those assets in an autonomous "co-scientist" platform.First, we introduce stereoelectronics-infused molecular graphs (SIMGs)-hypergraphs that encode lone pairs, bond orbitals and orbital interactions extracted from Natural Bond Orbital (NBO) analysis. A two-stage graph-neural pipeline approximates SIMGs, delivering high accuracy while remaining tractable for large molecules.Second, the work tackles biocatalysis, where enzyme-substrate activity is notoriously hard to predict. Leveraging ranking models, we build a recommender system that prioritizes enzymes for novel substrates.Third, we address data scarcity by constructing the largest experimental reaction dataset to date. High-resolution flow-injection MS, robotically prepared imine libraries, and reaction-based multiplexing yield 50,000 sample-compound pairs, capturing full kinetic and compositional profiles indispensable for next-generation ML models.Finally, these components integrate into Coscientist, an LLM-orchestrated agent that plans, executes and interprets experiments across heterogeneous lab hardware. Validation spans autonomous optimization of Pd-catalyzed couplings and multi-tool instrument control, illustrating how the platform can compress the design-make-test-analyze cycle.Collectively, the thesis advances molecular representation, dataset scale, and laboratory autonomy, laying a foundation for ML-driven experimentation in chemistry.
■590 ▼aSchool code: 0041.
■650 4▼aChemical engineering
■650 4▼aMaterials science
■650 4▼aMolecular chemistry
■653 ▼aStereoelectronics-infused molecular graphs
■653 ▼aBiocatalysis
■690 ▼a0542
■690 ▼a0431
■690 ▼a0794
■71020▼aCarnegie Mellon University▼bChemical Engineering.
■7730 ▼tDissertations Abstracts International▼g87-02B.
■790 ▼a0041
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17359009▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


