서브메뉴
검색
Theory of Learning in Neural Networks with Small Weight Perturbations
Theory of Learning in Neural Networks with Small Weight Perturbations
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211151330
- ISBN
- 9798382775937
- DDC
- 616
- 저자명
- Shan, Haozhe.
- 서명/저자
- Theory of Learning in Neural Networks with Small Weight Perturbations
- 발행사항
- [Sl] : Harvard University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 146 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 85-12, Section: B.
- 주기사항
- Advisor: Sompolinsky, Haim.
- 학위논문주기
- Thesis (Ph.D.)--Harvard University, 2024.
- 초록/해제
- 요약Learning multiple tasks in a non-stationary world requires continual learning (CL) -- the ability to accumulate and refine knowledge and skills over time. In neural networks (NN), realizing CL requires balancing stability (retaining benefits of previous learning in weights), and plasticity (efficient acquisition of new information). While CL occurs mundanely for biological NNs in the brain, artificial NNs in machine learning (ML) often fail catastrophically. This contrast poses two questions: (1) how does the brain handle the stability-plasticity dilemma and realize CL? (2) what causes CL to fail in artificial NNs and what could be done to rescue it? Towards answering these questions, this work presents a theoretical treatment of CL in NNs equipped with a simple, analytically tractable mechanism -- a weight-perturbation penalty that constrains the learning process to make small perturbations to weights.We first tested how the need to reduce learning-induced perturbations can explain neural mechanisms behind perceptual learning (PL) -- a well-studied experimental paradigm where animals exhibit long-lasting improvement in perceptual tasks following extensive training. While PL-induced physiological changes in sensory cortical areas are well documented, normative and mechanistic explanations of them are lacking. We hypothesized that the criticality of these areas for a broad range downstream tasks gives stability paramount importance. Thus, such areas should be modified with minimum perturbations (MP). To study its implications, we modeled the sensory hierarchy as a deep NN and developed a mean-field theory of the network in the limit of a large number of neurons and large number of examples. Our theory suggests that the input-output function of the network can be exactly mapped to that of a deep linear network, allowing us to characterize the space of solutions for the task as well as the MP solution within it. Interestingly, MP plasticity induces changes to weights and neural representations in all layers of the network, except for the readout weight vector. While weight changes in higher layers are not necessary for learning, they help reduce overall perturbation to the network. MP plasticity predicts physiological and behavioral changes that are largely consistent with experimental observations, suggesting MP as one of the potential learning principles in sensory areas in the adult brain.Generalizing beyond the setting of PL, we then used tools from statistical physics to develop a comprehensive theory of deep NNs learning sequences of arbitrary tasks with small weight perturbations. Our analytical results exactly describe how the input-output mapping of the network evolves as more tasks are learned sequentially. The degree of forgetting and transfer during CL is theoretically connected to relations between tasks, the network's architecture, and hyperparameters of the learning process. Of note, the theory identifies two scalar order parameters (OP) that succinctly capture input and rule similarity between tasks and suggests that they play related but diverging roles in determining CL outcomes. These OPs, directly computed from task data, are highly predictive of CL performance across a wide range of settings. The analysis also reveals how the architecture, including depth and whether there are task-dedicated readouts, strongly modulates the connection between task relations and CL performance. In particular, when the network contains task-dedicated readouts, our theory predicts three dramatically different CL regimes (or "phases"), determined by the task OPs and the amount of training data available. Sequentially learning tasks that are too dissimilar, as measured by the OPs, can lead to the surprising phenomenon of "catastrophic anterograde interference", where the network reaches zero training error on the new task but fails to generalize. Our results provide a rigorous treatment of the rich phenomena of CL in deep NNs and distinguish critical factors that promote or hinder CL.In conclusion, this work presents a theoretical analysis of how the need of stability-plasticity balance shapes learning in NNs. We hope that the results lay groundwork for further insights into neural mechanisms underlying CL in the brain as well as inspire practical algorithms for CL in artificial intelligence systems.
- 일반주제명
- Neurosciences
- 일반주제명
- Bioinformatics
- 키워드
- Neural networks
- 키워드
- Machine learning
- 기타저자
- Harvard University Medical Sciences
- 기본자료저록
- Dissertations Abstracts International. 85-12B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017161249
■00520250211151330
■006m o d
■007cr#unu||||||||
■020 ▼a9798382775937
■035 ▼a(MiAaPQ)AAI31240915
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a616
■1001 ▼aShan, Haozhe.▼0(orcid)0000-0002-1168-4861
■24510▼aTheory of Learning in Neural Networks with Small Weight Perturbations
■260 ▼a[Sl]▼bHarvard University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a146 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 85-12, Section: B.
■500 ▼aAdvisor: Sompolinsky, Haim.
■5021 ▼aThesis (Ph.D.)--Harvard University, 2024.
■520 ▼aLearning multiple tasks in a non-stationary world requires continual learning (CL) -- the ability to accumulate and refine knowledge and skills over time. In neural networks (NN), realizing CL requires balancing stability (retaining benefits of previous learning in weights), and plasticity (efficient acquisition of new information). While CL occurs mundanely for biological NNs in the brain, artificial NNs in machine learning (ML) often fail catastrophically. This contrast poses two questions: (1) how does the brain handle the stability-plasticity dilemma and realize CL? (2) what causes CL to fail in artificial NNs and what could be done to rescue it? Towards answering these questions, this work presents a theoretical treatment of CL in NNs equipped with a simple, analytically tractable mechanism -- a weight-perturbation penalty that constrains the learning process to make small perturbations to weights.We first tested how the need to reduce learning-induced perturbations can explain neural mechanisms behind perceptual learning (PL) -- a well-studied experimental paradigm where animals exhibit long-lasting improvement in perceptual tasks following extensive training. While PL-induced physiological changes in sensory cortical areas are well documented, normative and mechanistic explanations of them are lacking. We hypothesized that the criticality of these areas for a broad range downstream tasks gives stability paramount importance. Thus, such areas should be modified with minimum perturbations (MP). To study its implications, we modeled the sensory hierarchy as a deep NN and developed a mean-field theory of the network in the limit of a large number of neurons and large number of examples. Our theory suggests that the input-output function of the network can be exactly mapped to that of a deep linear network, allowing us to characterize the space of solutions for the task as well as the MP solution within it. Interestingly, MP plasticity induces changes to weights and neural representations in all layers of the network, except for the readout weight vector. While weight changes in higher layers are not necessary for learning, they help reduce overall perturbation to the network. MP plasticity predicts physiological and behavioral changes that are largely consistent with experimental observations, suggesting MP as one of the potential learning principles in sensory areas in the adult brain.Generalizing beyond the setting of PL, we then used tools from statistical physics to develop a comprehensive theory of deep NNs learning sequences of arbitrary tasks with small weight perturbations. Our analytical results exactly describe how the input-output mapping of the network evolves as more tasks are learned sequentially. The degree of forgetting and transfer during CL is theoretically connected to relations between tasks, the network's architecture, and hyperparameters of the learning process. Of note, the theory identifies two scalar order parameters (OP) that succinctly capture input and rule similarity between tasks and suggests that they play related but diverging roles in determining CL outcomes. These OPs, directly computed from task data, are highly predictive of CL performance across a wide range of settings. The analysis also reveals how the architecture, including depth and whether there are task-dedicated readouts, strongly modulates the connection between task relations and CL performance. In particular, when the network contains task-dedicated readouts, our theory predicts three dramatically different CL regimes (or "phases"), determined by the task OPs and the amount of training data available. Sequentially learning tasks that are too dissimilar, as measured by the OPs, can lead to the surprising phenomenon of "catastrophic anterograde interference", where the network reaches zero training error on the new task but fails to generalize. Our results provide a rigorous treatment of the rich phenomena of CL in deep NNs and distinguish critical factors that promote or hinder CL.In conclusion, this work presents a theoretical analysis of how the need of stability-plasticity balance shapes learning in NNs. We hope that the results lay groundwork for further insights into neural mechanisms underlying CL in the brain as well as inspire practical algorithms for CL in artificial intelligence systems.
■590 ▼aSchool code: 0084.
■650 4▼aNeurosciences
■650 4▼aBioinformatics
■653 ▼aNeural networks
■653 ▼aContinual learning
■653 ▼aMachine learning
■653 ▼aPerceptual learning
■653 ▼aMinimum perturbations
■690 ▼a0317
■690 ▼a0800
■690 ▼a0715
■71020▼aHarvard University▼bMedical Sciences.
■7730 ▼tDissertations Abstracts International▼g85-12B.
■790 ▼a0084
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17161249▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


