본문

서브메뉴

Theory of Learning in Neural Networks with Small Weight Perturbations
Theory of Learning in Neural Networks with Small Weight Perturbations
Theory of Learning in Neural Networks with Small Weight Perturbations

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211151330
ISBN  
9798382775937
DDC  
616
저자명  
Shan, Haozhe.
서명/저자  
Theory of Learning in Neural Networks with Small Weight Perturbations
발행사항  
[Sl] : Harvard University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
146 p
주기사항  
Source: Dissertations Abstracts International, Volume: 85-12, Section: B.
주기사항  
Advisor: Sompolinsky, Haim.
학위논문주기  
Thesis (Ph.D.)--Harvard University, 2024.
초록/해제  
요약Learning multiple tasks in a non-stationary world requires continual learning (CL) -- the ability to accumulate and refine knowledge and skills over time. In neural networks (NN), realizing CL requires balancing stability (retaining benefits of previous learning in weights), and plasticity (efficient acquisition of new information). While CL occurs mundanely for biological NNs in the brain, artificial NNs in machine learning (ML) often fail catastrophically. This contrast poses two questions: (1) how does the brain handle the stability-plasticity dilemma and realize CL? (2) what causes CL to fail in artificial NNs and what could be done to rescue it? Towards answering these questions, this work presents a theoretical treatment of CL in NNs equipped with a simple, analytically tractable mechanism -- a weight-perturbation penalty that constrains the learning process to make small perturbations to weights.We first tested how the need to reduce learning-induced perturbations can explain neural mechanisms behind perceptual learning (PL) -- a well-studied experimental paradigm where animals exhibit long-lasting improvement in perceptual tasks following extensive training. While PL-induced physiological changes in sensory cortical areas are well documented, normative and mechanistic explanations of them are lacking. We hypothesized that the criticality of these areas for a broad range downstream tasks gives stability paramount importance. Thus, such areas should be modified with minimum perturbations (MP). To study its implications, we modeled the sensory hierarchy as a deep NN and developed a mean-field theory of the network in the limit of a large number of neurons and large number of examples. Our theory suggests that the input-output function of the network can be exactly mapped to that of a deep linear network, allowing us to characterize the space of solutions for the task as well as the MP solution within it. Interestingly, MP plasticity induces changes to weights and neural representations in all layers of the network, except for the readout weight vector. While weight changes in higher layers are not necessary for learning, they help reduce overall perturbation to the network. MP plasticity predicts physiological and behavioral changes that are largely consistent with experimental observations, suggesting MP as one of the potential learning principles in sensory areas in the adult brain.Generalizing beyond the setting of PL, we then used tools from statistical physics to develop a comprehensive theory of deep NNs learning sequences of arbitrary tasks with small weight perturbations. Our analytical results exactly describe how the input-output mapping of the network evolves as more tasks are learned sequentially. The degree of forgetting and transfer during CL is theoretically connected to relations between tasks, the network's architecture, and hyperparameters of the learning process. Of note, the theory identifies two scalar order parameters (OP) that succinctly capture input and rule similarity between tasks and suggests that they play related but diverging roles in determining CL outcomes. These OPs, directly computed from task data, are highly predictive of CL performance across a wide range of settings. The analysis also reveals how the architecture, including depth and whether there are task-dedicated readouts, strongly modulates the connection between task relations and CL performance. In particular, when the network contains task-dedicated readouts, our theory predicts three dramatically different CL regimes (or "phases"), determined by the task OPs and the amount of training data available. Sequentially learning tasks that are too dissimilar, as measured by the OPs, can lead to the surprising phenomenon of "catastrophic anterograde interference", where the network reaches zero training error on the new task but fails to generalize. Our results provide a rigorous treatment of the rich phenomena of CL in deep NNs and distinguish critical factors that promote or hinder CL.In conclusion, this work presents a theoretical analysis of how the need of stability-plasticity balance shapes learning in NNs. We hope that the results lay groundwork for further insights into neural mechanisms underlying CL in the brain as well as inspire practical algorithms for CL in artificial intelligence systems.
일반주제명  
Neurosciences
일반주제명  
Bioinformatics
키워드  
Neural networks
키워드  
Continual learning
키워드  
Machine learning
키워드  
Perceptual learning
키워드  
Minimum perturbations
기타저자  
Harvard University Medical Sciences
기본자료저록  
Dissertations Abstracts International. 85-12B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017161249
■00520250211151330
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798382775937
■035    ▼a(MiAaPQ)AAI31240915
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a616
■1001  ▼aShan,  Haozhe.▼0(orcid)0000-0002-1168-4861
■24510▼aTheory  of  Learning  in  Neural  Networks  with  Small  Weight  Perturbations
■260    ▼a[Sl]▼bHarvard  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a146  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  85-12,  Section:  B.
■500    ▼aAdvisor:  Sompolinsky,  Haim.
■5021  ▼aThesis  (Ph.D.)--Harvard  University,  2024.
■520    ▼aLearning  multiple  tasks  in  a  non-stationary  world  requires  continual  learning  (CL)  --  the  ability  to  accumulate  and  refine  knowledge  and  skills  over  time.  In  neural  networks  (NN),  realizing  CL  requires  balancing  stability  (retaining  benefits  of  previous  learning  in  weights),  and  plasticity  (efficient  acquisition  of  new  information).  While  CL  occurs  mundanely  for  biological  NNs  in  the  brain,  artificial  NNs  in  machine  learning  (ML)  often  fail  catastrophically.  This  contrast  poses  two  questions:  (1)  how  does  the  brain  handle  the  stability-plasticity  dilemma  and  realize  CL?  (2)  what  causes  CL  to  fail  in  artificial  NNs  and  what  could  be  done  to  rescue  it?  Towards  answering  these  questions,  this  work  presents  a  theoretical  treatment  of  CL  in  NNs  equipped  with  a  simple,  analytically  tractable  mechanism  --  a  weight-perturbation  penalty  that  constrains  the  learning  process  to  make  small  perturbations  to  weights.We  first  tested  how  the  need  to  reduce  learning-induced  perturbations  can  explain  neural  mechanisms  behind  perceptual  learning  (PL)  --  a  well-studied  experimental  paradigm  where  animals  exhibit  long-lasting  improvement  in  perceptual  tasks  following  extensive  training.  While  PL-induced  physiological  changes  in  sensory  cortical  areas  are  well  documented,  normative  and  mechanistic  explanations  of  them  are  lacking.  We  hypothesized  that  the  criticality  of  these  areas  for  a  broad  range  downstream  tasks  gives  stability  paramount  importance.  Thus,  such  areas  should  be  modified  with  minimum  perturbations  (MP).  To  study  its  implications,  we  modeled  the  sensory  hierarchy  as  a  deep  NN  and  developed  a  mean-field  theory  of  the  network  in  the  limit  of  a  large  number  of  neurons  and  large  number  of  examples.  Our  theory  suggests  that  the  input-output  function  of  the  network  can  be  exactly  mapped  to  that  of  a  deep  linear  network,  allowing  us  to  characterize  the  space  of  solutions  for  the  task  as  well  as  the  MP  solution  within  it.  Interestingly,  MP  plasticity  induces  changes  to  weights  and  neural  representations  in  all  layers  of  the  network,  except  for  the  readout  weight  vector.  While  weight  changes  in  higher  layers  are  not  necessary  for  learning,  they  help  reduce  overall  perturbation  to  the  network.  MP  plasticity  predicts  physiological  and  behavioral  changes  that  are  largely  consistent  with  experimental  observations,  suggesting  MP  as  one  of  the  potential  learning  principles  in  sensory  areas  in  the  adult  brain.Generalizing  beyond  the  setting  of  PL,  we  then  used  tools  from  statistical  physics  to  develop  a  comprehensive  theory  of  deep  NNs  learning  sequences  of  arbitrary  tasks  with  small  weight  perturbations.  Our  analytical  results  exactly  describe  how  the  input-output  mapping  of  the  network  evolves  as  more  tasks  are  learned  sequentially.  The  degree  of  forgetting  and  transfer  during  CL  is  theoretically  connected  to  relations  between  tasks,  the  network's  architecture,  and  hyperparameters  of  the  learning  process.  Of  note,  the  theory  identifies  two  scalar  order  parameters  (OP)  that  succinctly  capture  input  and  rule  similarity  between  tasks  and  suggests  that  they  play  related  but  diverging  roles  in  determining  CL  outcomes.  These  OPs,  directly  computed  from  task  data,  are  highly  predictive  of  CL  performance  across  a  wide  range  of  settings.  The  analysis  also  reveals  how  the  architecture,  including  depth  and  whether  there  are  task-dedicated  readouts,  strongly  modulates  the  connection  between  task  relations  and  CL  performance.  In  particular,  when  the  network  contains  task-dedicated  readouts,  our  theory  predicts  three  dramatically  different  CL  regimes  (or  "phases"),  determined  by  the  task  OPs  and  the  amount  of  training  data  available.  Sequentially  learning  tasks  that  are  too  dissimilar,  as  measured  by  the  OPs,  can  lead  to  the  surprising  phenomenon  of  "catastrophic  anterograde  interference",  where  the  network  reaches  zero  training  error  on  the  new  task  but  fails  to  generalize.  Our  results  provide  a  rigorous  treatment  of  the  rich  phenomena  of  CL  in  deep  NNs  and  distinguish  critical  factors  that  promote  or  hinder  CL.In  conclusion,  this  work  presents  a  theoretical  analysis  of  how  the  need  of  stability-plasticity  balance  shapes  learning  in  NNs.  We  hope  that  the  results  lay  groundwork  for  further  insights  into  neural  mechanisms  underlying  CL  in  the  brain  as  well  as  inspire  practical  algorithms  for  CL  in  artificial  intelligence  systems.
■590    ▼aSchool  code:  0084.
■650  4▼aNeurosciences
■650  4▼aBioinformatics
■653    ▼aNeural  networks
■653    ▼aContinual  learning
■653    ▼aMachine  learning
■653    ▼aPerceptual  learning
■653    ▼aMinimum  perturbations
■690    ▼a0317
■690    ▼a0800
■690    ▼a0715
■71020▼aHarvard  University▼bMedical  Sciences.
■7730  ▼tDissertations  Abstracts  International▼g85-12B.
■790    ▼a0084
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17161249▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF12322 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.