서브메뉴
검색
Modelling Sequence and Structure Towards Functional Protein Design
Modelling Sequence and Structure Towards Functional Protein Design
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211152823
- ISBN
- 9798346567646
- DDC
- 574
- 서명/저자
- Modelling Sequence and Structure Towards Functional Protein Design
- 발행사항
- [Sl] : Harvard University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 210 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-05, Section: B.
- 주기사항
- Advisor: Marsk, Debora S.
- 학위논문주기
- Thesis (Ph.D.)--Harvard University, 2024.
- 초록/해제
- 요약Millenia of evolutionary experiments have produced an extensive universe of natural macromolecular machines - proteins - that perform the variety of complex functions needed to make up a cell. In the last few decades, advances in protein engineering technologies, including the adoption of machine learning methods, have enabled us to bend and reform nature's designs towards our human needs. The advent of generative machine learning models trained on evolutionary data has enabled us to leverage nature's experiments along with years of domain knowledge to significantly move the needle on what kinds of proteins and functions we can possibly design. While these models have massive promise, there is still much to understand about i) what these models are learning, ii) how well they are learning, and iii) which models are useful for which design task. This thesis provides tools and insights to the field to shed light on these questions and thus advance our ability to engineer proteins for the functions we want.We begin with the need to identify what models perform better than others and what biological design tasks they may be useful for. Towards this, chapter 1 details our curation of the largest benchmarking dataset for generative protein models for fitness prediction, which we used to identify functional advantages for particular classes of generative models. This evaluation paradigm relies on functional measurements, which may not be available for any given protein an engineer is interested in. Thus, in chapter 2 we develop novel, statistically motivated kernel-based evaluation metrics that can be used to verify how accurately and reliably a conditional generative model has learned the distribution of the protein of interest; this provides a practitioner with helpful information about how well their model might perform for their task a priori. For highly complex functions in highly local sequence space, we argue that focused experimental data are needed to get engineering gains. In chapter 3 we discuss how machine learning models can improve the efficiency of experimental pipelines and increase our design capabilities, with a case-study on machine learning-assisted antibody optimization.
- 일반주제명
- Bioinformatics
- 일반주제명
- Applied mathematics
- 일반주제명
- Bioengineering
- 일반주제명
- Evolution & development
- 키워드
- Generative AI
- 키워드
- Machine learning
- 키워드
- Millenia
- 키워드
- Protein design
- 기타저자
- Harvard University Medical Sciences
- 기본자료저록
- Dissertations Abstracts International. 86-05B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017164029
■00520250211152823
■006m o d
■007cr#unu||||||||
■020 ▼a9798346567646
■035 ▼a(MiAaPQ)AAI31559752
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a574
■1001 ▼aPaul, Steffanie B.▼0(orcid)0000-0001-7306-4863
■24510▼aModelling Sequence and Structure Towards Functional Protein Design
■260 ▼a[Sl]▼bHarvard University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a210 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-05, Section: B.
■500 ▼aAdvisor: Marsk, Debora S.
■5021 ▼aThesis (Ph.D.)--Harvard University, 2024.
■520 ▼aMillenia of evolutionary experiments have produced an extensive universe of natural macromolecular machines - proteins - that perform the variety of complex functions needed to make up a cell. In the last few decades, advances in protein engineering technologies, including the adoption of machine learning methods, have enabled us to bend and reform nature's designs towards our human needs. The advent of generative machine learning models trained on evolutionary data has enabled us to leverage nature's experiments along with years of domain knowledge to significantly move the needle on what kinds of proteins and functions we can possibly design. While these models have massive promise, there is still much to understand about i) what these models are learning, ii) how well they are learning, and iii) which models are useful for which design task. This thesis provides tools and insights to the field to shed light on these questions and thus advance our ability to engineer proteins for the functions we want.We begin with the need to identify what models perform better than others and what biological design tasks they may be useful for. Towards this, chapter 1 details our curation of the largest benchmarking dataset for generative protein models for fitness prediction, which we used to identify functional advantages for particular classes of generative models. This evaluation paradigm relies on functional measurements, which may not be available for any given protein an engineer is interested in. Thus, in chapter 2 we develop novel, statistically motivated kernel-based evaluation metrics that can be used to verify how accurately and reliably a conditional generative model has learned the distribution of the protein of interest; this provides a practitioner with helpful information about how well their model might perform for their task a priori. For highly complex functions in highly local sequence space, we argue that focused experimental data are needed to get engineering gains. In chapter 3 we discuss how machine learning models can improve the efficiency of experimental pipelines and increase our design capabilities, with a case-study on machine learning-assisted antibody optimization.
■590 ▼aSchool code: 0084.
■650 4▼aBioinformatics
■650 4▼aApplied mathematics
■650 4▼aBioengineering
■650 4▼aEvolution & development
■653 ▼aGenerative AI
■653 ▼aMachine learning
■653 ▼aProtein engineering
■653 ▼aMillenia
■653 ▼aProtein design
■690 ▼a0715
■690 ▼a0364
■690 ▼a0202
■690 ▼a0412
■71020▼aHarvard University▼bMedical Sciences.
■7730 ▼tDissertations Abstracts International▼g86-05B.
■790 ▼a0084
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164029▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


