서브메뉴
검색
Evaluation of the Generalizability of Machine Learning-Assisted Protein Engineering Methods
Evaluation of the Generalizability of Machine Learning-Assisted Protein Engineering Methods
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202104748
- ISBN
- 9798290652580
- DDC
- 400
- 서명/저자
- Evaluation of the Generalizability of Machine Learning-Assisted Protein Engineering Methods
- 발행사항
- [Sl] : California Institute of Technology, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 230 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-01, Section: B.
- 주기사항
- Advisor: Arnold, Frances Hamilton;Yue, Yisong.
- 학위논문주기
- Thesis (Ph.D.)--California Institute of Technology, 2025.
- 초록/해제
- 요약Engineered proteins can carry out a vast array of functions and have become indispensable across numerous industrial applications. To accelerate wet-lab protein engineering efforts, machine learning-based methods have advanced rapidly. However, a gap remains between state-of-the-art machine learning methods and their practical adoption. A key factor contributing to this disconnect is the lack of application-relevant benchmarking and generalizable insights across protein engineering tasks. This thesis evaluates machine learning-assisted protein engineering approaches to identify generalizable strategies. The central problem considered is learning the mapping from protein sequence to function-known as the fitness landscape-to enable the prediction of unseen variant fitness. Chapter 1introduces the background and context for machine learning-assisted protein engineering and highlights the practical constraint of limited experimental budgets. Chapter 2investigates transfer learning, which leverages models pretrained on large protein sequence databases to generate informative representations for modeling task specific sequence-function relationships. Evaluation across ten diverse tasks shows that while transfer learning is effective in structure prediction, it underperforms in variant fitness prediction-a key objective in protein engineering. Chapter 3evaluates alternative strategies with a focus on combinatorial fitness landscapes, a common setting in protein engineering. Across 16 diverse landscapes, focused trainingimproves the performance of various machine learning approaches by strategically selecting training variants using zero-shot predictors, which estimate variant fitness from auxiliary information without relying on experimental data. Building on these insights, Chapter 4addresses the specific challenge of engineering enzymes-proteins that convert substrates into products-for novel chemistries. While six general zero-shot predictors without substrate information can predict enzyme activity on non-native substrates, they fail on more out-of-distribution, new-to-naturechemistries. Incorporating substrate information into zero-shot predictors leads to more generalizable performance across all tested chemistries, spanning 22 substrates. Overall, this thesis identifies generalizable strategies for machine learning-assisted protein engineering by systematically evaluating and improving how sequence-to-function relationships are modeled across diverse tasks.
- 일반주제명
- Language
- 일반주제명
- Dihydrofolate reductase
- 일반주제명
- Software
- 일반주제명
- Bioengineering
- 일반주제명
- Neural networks
- 일반주제명
- Community
- 일반주제명
- Engineering
- 일반주제명
- Libraries
- 일반주제명
- Mutagenesis
- 기타저자
- California Institute of Technology Biology and Biological Engineering
- 기본자료저록
- Dissertations Abstracts International. 87-01B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017358762
■00520260202104748
■006m o d
■007cr#unu||||||||
■020 ▼a9798290652580
■035 ▼a(MiAaPQ)AAI32151309
■035 ▼a(MiAaPQ)Caltech17182
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a400
■1001 ▼aLi, Francesca-Zhoufan.
■24510▼aEvaluation of the Generalizability of Machine Learning-Assisted Protein Engineering Methods
■260 ▼a[Sl]▼bCalifornia Institute of Technology▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a230 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-01, Section: B.
■500 ▼aAdvisor: Arnold, Frances Hamilton;Yue, Yisong.
■5021 ▼aThesis (Ph.D.)--California Institute of Technology, 2025.
■520 ▼aEngineered proteins can carry out a vast array of functions and have become indispensable across numerous industrial applications. To accelerate wet-lab protein engineering efforts, machine learning-based methods have advanced rapidly. However, a gap remains between state-of-the-art machine learning methods and their practical adoption. A key factor contributing to this disconnect is the lack of application-relevant benchmarking and generalizable insights across protein engineering tasks. This thesis evaluates machine learning-assisted protein engineering approaches to identify generalizable strategies. The central problem considered is learning the mapping from protein sequence to function-known as the fitness landscape-to enable the prediction of unseen variant fitness. Chapter 1introduces the background and context for machine learning-assisted protein engineering and highlights the practical constraint of limited experimental budgets. Chapter 2investigates transfer learning, which leverages models pretrained on large protein sequence databases to generate informative representations for modeling task specific sequence-function relationships. Evaluation across ten diverse tasks shows that while transfer learning is effective in structure prediction, it underperforms in variant fitness prediction-a key objective in protein engineering. Chapter 3evaluates alternative strategies with a focus on combinatorial fitness landscapes, a common setting in protein engineering. Across 16 diverse landscapes, focused trainingimproves the performance of various machine learning approaches by strategically selecting training variants using zero-shot predictors, which estimate variant fitness from auxiliary information without relying on experimental data. Building on these insights, Chapter 4addresses the specific challenge of engineering enzymes-proteins that convert substrates into products-for novel chemistries. While six general zero-shot predictors without substrate information can predict enzyme activity on non-native substrates, they fail on more out-of-distribution, new-to-naturechemistries. Incorporating substrate information into zero-shot predictors leads to more generalizable performance across all tested chemistries, spanning 22 substrates. Overall, this thesis identifies generalizable strategies for machine learning-assisted protein engineering by systematically evaluating and improving how sequence-to-function relationships are modeled across diverse tasks.
■590 ▼aSchool code: 0037.
■650 4▼aLanguage
■650 4▼aDihydrofolate reductase
■650 4▼aSoftware
■650 4▼aBioengineering
■650 4▼aNeural networks
■650 4▼aCommunity
■650 4▼aEngineering
■650 4▼aLibraries
■650 4▼aMutagenesis
■690 ▼a0202
■690 ▼a0800
■690 ▼a0679
■690 ▼a0537
■71020▼aCalifornia Institute of Technology▼bBiology and Biological Engineering.
■7730 ▼tDissertations Abstracts International▼g87-01B.
■790 ▼a0037
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17358762▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


