본문

서브메뉴

Modelling Sequence and Structure Towards Functional Protein Design
Modelling Sequence and Structure Towards Functional Protein Design
Modelling Sequence and Structure Towards Functional Protein Design

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211152823
ISBN  
9798346567646
DDC  
574
저자명  
Paul, Steffanie B.
서명/저자  
Modelling Sequence and Structure Towards Functional Protein Design
발행사항  
[Sl] : Harvard University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
210 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-05, Section: B.
주기사항  
Advisor: Marsk, Debora S.
학위논문주기  
Thesis (Ph.D.)--Harvard University, 2024.
초록/해제  
요약Millenia of evolutionary experiments have produced an extensive universe of natural macromolecular machines - proteins - that perform the variety of complex functions needed to make up a cell. In the last few decades, advances in protein engineering technologies, including the adoption of machine learning methods, have enabled us to bend and reform nature's designs towards our human needs. The advent of generative machine learning models trained on evolutionary data has enabled us to leverage nature's experiments along with years of domain knowledge to significantly move the needle on what kinds of proteins and functions we can possibly design. While these models have massive promise, there is still much to understand about i) what these models are learning, ii) how well they are learning, and iii) which models are useful for which design task. This thesis provides tools and insights to the field to shed light on these questions and thus advance our ability to engineer proteins for the functions we want.We begin with the need to identify what models perform better than others and what biological design tasks they may be useful for. Towards this, chapter 1 details our curation of the largest benchmarking dataset for generative protein models for fitness prediction, which we used to identify functional advantages for particular classes of generative models. This evaluation paradigm relies on functional measurements, which may not be available for any given protein an engineer is interested in. Thus, in chapter 2 we develop novel, statistically motivated kernel-based evaluation metrics that can be used to verify how accurately and reliably a conditional generative model has learned the distribution of the protein of interest; this provides a practitioner with helpful information about how well their model might perform for their task a priori. For highly complex functions in highly local sequence space, we argue that focused experimental data are needed to get engineering gains. In chapter 3 we discuss how machine learning models can improve the efficiency of experimental pipelines and increase our design capabilities, with a case-study on machine learning-assisted antibody optimization.
일반주제명  
Bioinformatics
일반주제명  
Applied mathematics
일반주제명  
Bioengineering
일반주제명  
Evolution & development
키워드  
Generative AI
키워드  
Machine learning
키워드  
Protein engineering
키워드  
Millenia
키워드  
Protein design
기타저자  
Harvard University Medical Sciences
기본자료저록  
Dissertations Abstracts International. 86-05B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017164029
■00520250211152823
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798346567646
■035    ▼a(MiAaPQ)AAI31559752
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a574
■1001  ▼aPaul,  Steffanie  B.▼0(orcid)0000-0001-7306-4863
■24510▼aModelling  Sequence  and  Structure  Towards  Functional  Protein  Design
■260    ▼a[Sl]▼bHarvard  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a210  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-05,  Section:  B.
■500    ▼aAdvisor:  Marsk,  Debora  S.
■5021  ▼aThesis  (Ph.D.)--Harvard  University,  2024.
■520    ▼aMillenia  of  evolutionary  experiments  have  produced  an  extensive  universe  of  natural  macromolecular  machines  -  proteins  -  that  perform  the  variety  of  complex  functions  needed  to  make  up  a  cell.  In  the  last  few  decades,  advances  in  protein  engineering  technologies,  including  the  adoption  of  machine  learning  methods,  have  enabled  us  to  bend  and  reform  nature's  designs  towards  our  human  needs.  The  advent  of  generative  machine  learning  models  trained  on  evolutionary  data  has  enabled  us  to  leverage  nature's  experiments  along  with  years  of  domain  knowledge  to  significantly  move  the  needle  on  what  kinds  of  proteins  and  functions  we  can  possibly  design.  While  these  models  have  massive  promise,  there  is  still  much  to  understand  about  i)  what  these  models  are  learning,  ii)  how  well  they  are  learning,  and  iii)  which  models  are  useful  for  which  design  task.  This  thesis  provides  tools  and  insights  to  the  field  to  shed  light  on  these  questions  and  thus  advance  our  ability  to  engineer  proteins  for  the  functions  we  want.We  begin  with  the  need  to  identify  what  models  perform  better  than  others  and  what  biological  design  tasks  they  may  be  useful  for.  Towards  this,  chapter  1  details  our  curation  of  the  largest  benchmarking  dataset  for  generative  protein  models  for  fitness  prediction,  which  we  used  to  identify  functional  advantages  for  particular  classes  of  generative  models.  This  evaluation  paradigm  relies  on  functional  measurements,  which  may  not  be  available  for  any  given  protein  an  engineer  is  interested  in.  Thus,  in  chapter  2  we  develop  novel,  statistically  motivated  kernel-based  evaluation  metrics  that  can  be  used  to  verify  how  accurately  and  reliably  a  conditional  generative  model  has  learned  the  distribution  of  the  protein  of  interest;  this  provides  a  practitioner  with  helpful  information  about  how  well  their  model  might  perform  for  their  task  a  priori.  For  highly  complex  functions  in  highly  local  sequence  space,  we  argue  that  focused  experimental  data  are  needed  to  get  engineering  gains.  In  chapter  3  we  discuss  how  machine  learning  models  can  improve  the  efficiency  of  experimental  pipelines  and  increase  our  design  capabilities,  with  a  case-study  on  machine  learning-assisted  antibody  optimization.
■590    ▼aSchool  code:  0084.
■650  4▼aBioinformatics
■650  4▼aApplied  mathematics
■650  4▼aBioengineering
■650  4▼aEvolution  &  development
■653    ▼aGenerative  AI
■653    ▼aMachine  learning
■653    ▼aProtein  engineering
■653    ▼aMillenia
■653    ▼aProtein  design
■690    ▼a0715
■690    ▼a0364
■690    ▼a0202
■690    ▼a0412
■71020▼aHarvard  University▼bMedical  Sciences.
■7730  ▼tDissertations  Abstracts  International▼g86-05B.
■790    ▼a0084
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164029▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF11835 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.