본문

서브메뉴

Building Reliable Machine Learning Systems for Neuroscience
Building Reliable Machine Learning Systems for Neuroscience
Building Reliable Machine Learning Systems for Neuroscience

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211151342
ISBN  
9798382787374
DDC  
616
저자명  
Buchanan, E. Kelly.
서명/저자  
Building Reliable Machine Learning Systems for Neuroscience
발행사항  
[Sl] : Columbia University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
258 p
주기사항  
Source: Dissertations Abstracts International, Volume: 85-12, Section: B.
주기사항  
Advisor: Paninski, Liam;Cunningham, John P.
학위논문주기  
Thesis (Ph.D.)--Columbia University, 2024.
초록/해제  
요약Neuroscience as a field is collecting more data than at any other time in history. The scale of this data allows us to ask fundamental questions about the mechanisms of brain function, the basis of behavior, and the development of disorders. Our ambitious goals as well as the abundance of data being recorded call for reproducible, reliable, and accessible systems to push the field forward. While we have made great strides in building reproducible and accessible machine learning (ML) systems for neuroscience, reliability remains a major issue.In this dissertation, we show that we can leverage existing data and domain expert knowledge to build more reliable ML systems to study animal behavior. First, we consider animal pose estimation, a crucial component in many scientific investigations. Typical transfer learning ML methods for behavioral tracking treat each video frame and object to be tracked independently. We improve on this by leveraging the rich spatial and temporal structures pervasive in behavioral videos. Our resulting weakly supervised models achieve significantly more robust tracking. Our tools allow us to achieve improved results when we have imperfect, limited data while requiring users to label fewer training frames and speeding up training. We can more accurately process raw video data and learn interpretable units of behavior. In turn, these improvements enhance performance on downstream applications.Next, we consider a ubiquitous approach to (attempt to) improve the reliability of ML methods, namely combining the predictions of multiple models, also known as deep ensembling. Ensembles of classical ML predictors, such as random forests, improve metrics such as accuracy by well-understood mechanisms such as improving diversity. However, in the case of deep ensembles, there is an open methodological question as to whether, given the choice between a deep ensemble and a single neural network with similar accuracy, one model is truly preferable over the other. Via careful experiments across a range of benchmark datasets and deep learning models, we demonstrate limitations to the purported benefits of deep ensembles. Our results challenge common assumptions regarding the effectiveness of deep ensembles and the "diversity" principles underpinning their success, especially with regards to important metrics for reliability, such as out-of-distribution (OOD) performance and effective robustness. We conduct additional studies of the effects of using deep ensembles when certain groups in the dataset are underrepresented (so-called "long tail" data), a setting whose importance in neuroscience applications is revealed by our aforementioned work.Altogether, our results demonstrate the essential importance of both holistic systems work and fundamental methodological work to understand the best ways to apply the benefits of modern machine learning to the unique challenges of neuroscience data analysis pipelines. To conclude the dissertation, we outline challenges and opportunities in building next-generation ML systems.
일반주제명  
Neurosciences
일반주제명  
Computer science
키워드  
Animal pose estimation
키워드  
Behavioral action segmentation
키워드  
Deep ensembles
키워드  
Machine learning
키워드  
Reliable Deep Learning
기타저자  
Columbia University Neurobiology and Behavior
기본자료저록  
Dissertations Abstracts International. 85-12B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017161340
■00520250211151342
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798382787374
■035    ▼a(MiAaPQ)AAI31242323
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a616
■1001  ▼aBuchanan,  E.  Kelly.
■24510▼aBuilding  Reliable  Machine  Learning  Systems  for  Neuroscience
■260    ▼a[Sl]▼bColumbia  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a258  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  85-12,  Section:  B.
■500    ▼aAdvisor:  Paninski,  Liam;Cunningham,  John  P.
■5021  ▼aThesis  (Ph.D.)--Columbia  University,  2024.
■520    ▼aNeuroscience  as  a  field  is  collecting  more  data  than  at  any  other  time  in  history.  The  scale  of  this  data  allows  us  to  ask  fundamental  questions  about  the  mechanisms  of  brain  function,  the  basis  of  behavior,  and  the  development  of  disorders.  Our  ambitious  goals  as  well  as  the  abundance  of  data  being  recorded  call  for  reproducible,  reliable,  and  accessible  systems  to  push  the  field  forward.  While  we  have  made  great  strides  in  building  reproducible  and  accessible  machine  learning  (ML)  systems  for  neuroscience,  reliability  remains  a  major  issue.In  this  dissertation,  we  show  that  we  can  leverage  existing  data  and  domain  expert  knowledge  to  build  more  reliable  ML  systems  to  study  animal  behavior.  First,  we  consider  animal  pose  estimation,  a  crucial  component  in  many  scientific  investigations.  Typical  transfer  learning  ML  methods  for  behavioral  tracking  treat  each  video  frame  and  object  to  be  tracked  independently.  We  improve  on  this  by  leveraging  the  rich  spatial  and  temporal  structures  pervasive  in  behavioral  videos.  Our  resulting  weakly  supervised  models  achieve  significantly  more  robust  tracking.  Our  tools  allow  us  to  achieve  improved  results  when  we  have  imperfect,  limited  data  while  requiring  users  to  label  fewer  training  frames  and  speeding  up  training.  We  can  more  accurately  process  raw  video  data  and  learn  interpretable  units  of  behavior.  In  turn,  these  improvements  enhance  performance  on  downstream  applications.Next,  we  consider  a  ubiquitous  approach  to  (attempt  to)  improve  the  reliability  of  ML  methods,  namely  combining  the  predictions  of  multiple  models,  also  known  as  deep  ensembling.  Ensembles  of  classical  ML  predictors,  such  as  random  forests,  improve  metrics  such  as  accuracy  by  well-understood  mechanisms  such  as  improving  diversity.  However,  in  the  case  of  deep  ensembles,  there  is  an  open  methodological  question  as  to  whether,  given  the  choice  between  a  deep  ensemble  and  a  single  neural  network  with  similar  accuracy,  one  model  is  truly  preferable  over  the  other.  Via  careful  experiments  across  a  range  of  benchmark  datasets  and  deep  learning  models,  we  demonstrate  limitations  to  the  purported  benefits  of  deep  ensembles.  Our  results  challenge  common  assumptions  regarding  the  effectiveness  of  deep  ensembles  and  the  "diversity"  principles  underpinning  their  success,  especially  with  regards  to  important  metrics  for  reliability,  such  as  out-of-distribution  (OOD)  performance  and  effective  robustness.  We  conduct  additional  studies  of  the  effects  of  using  deep  ensembles  when  certain  groups  in  the  dataset  are  underrepresented  (so-called  "long  tail"  data),  a  setting  whose  importance  in  neuroscience  applications  is  revealed  by  our  aforementioned  work.Altogether,  our  results  demonstrate  the  essential  importance  of  both  holistic  systems  work  and  fundamental  methodological  work  to  understand  the  best  ways  to  apply  the  benefits  of  modern  machine  learning  to  the  unique  challenges  of  neuroscience  data  analysis  pipelines.  To  conclude  the  dissertation,  we  outline  challenges  and  opportunities  in  building  next-generation  ML  systems.
■590    ▼aSchool  code:  0054.
■650  4▼aNeurosciences
■650  4▼aComputer  science
■653    ▼aAnimal  pose  estimation
■653    ▼aBehavioral  action  segmentation
■653    ▼aDeep  ensembles
■653    ▼aMachine  learning
■653    ▼aReliable  Deep  Learning
■690    ▼a0317
■690    ▼a0984
■690    ▼a0800
■71020▼aColumbia  University▼bNeurobiology  and  Behavior.
■7730  ▼tDissertations  Abstracts  International▼g85-12B.
■790    ▼a0054
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17161340▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF09468 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.