서브메뉴
검색
Neural Network Models of Learning and Generalization
Neural Network Models of Learning and Generalization
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202104750
- ISBN
- 9798290657219
- DDC
- 330
- 서명/저자
- Neural Network Models of Learning and Generalization
- 발행사항
- [Sl] : California Institute of Technology, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 250 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-01, Section: B.
- 주기사항
- Advisor: Rangel, Antonio.
- 학위논문주기
- Thesis (Ph.D.)--California Institute of Technology, 2025.
- 초록/해제
- 요약Neural networks have emerged as powerful models for understanding both biological and artificial intelligence, yet significant questions remain about how these systems develop rich, generalizable representations of the world. This thesis investigates fundamental principles of learning and generalization across four interconnected domains, bridging insights from theoretical neuroscience and artificial intelligence to advance our understanding of intelligent systems.In Chapter I, we address a central question in associative learning: how do neural circuits learn to associate concepts with one another? We propose a recurrent neural network model incorporating two critical features of cortical architecture-mixed selectivity and compartmentalized neurons. These architectural inductive biases enable a biologically plausible learning rule that achieves stimulus substitution, where neurons respond identically to a conditioned stimulus as they would to the associated unconditioned stimulus. Our model explains a remarkable range of conditioning phenomena under conditions in which traditional associative models fail, highlighting how the cortical architecture may confer significant evolutionary advantages for flexible learning.Chapter II pivots from the static mappings between concepts learned in Chapter I to explore how neural systems develop the precise synaptic connectivity required to establish dynamic mappings for path integration-the ability to maintain an internal sense of direction without external cues. We demonstrate that the same principles of compartmentalized learning can shape networks that accurately track angular position in darkness. Applied to the Drosophilahead direction system, our model develops connectivity patterns strikingly similar to those observed experimentally, with continuous attractor (CAN) dynamics emerging naturally from learning. This offers a novel perspective on how precisely calibrated neural circuits can develop through experience, rather than requiring genetic pre-specification, and explains experimental findings where animals adapt their internal representation when sensory experience changes.In Chapter III, we establish a theoretical framework explaining how disentangled representations-internal models that isolate independent factors of variation in the world-emerge from multi-task learning. We prove that any system competent at multiple related tasks must implicitly represent the underlying latent variables in a linearly decodable form when sufficient tasks are learned. These theoretical guarantees align with experimental results showing neural networks develop generalizable representations when trained on multiple tasks simultaneously. This work reveals a fundamental connection between task diversity and representation quality, with implications for biological cognition and artificial intelligence design, particularly explaining why modern transformer models may develop human-interpretable concepts, and how brains may acquire their impressive zero-shot generalization ability.Chapter IV proposes leveraging Large Language Models as cognitive tools for evaluating latent factor hypotheses for psychology, leveraging the theoretical insights from Chapter III. It suggests that the self-consistency of an LLM's responses given hypothesized psychological factors could serve as a metric for hypothesis evaluation. While preliminary, this approach represents a novel computational methodology for psychology that could transform how hypotheses for human cognition are developed and refined.All chapters are supported by corresponding Appendices that go deeper in particular details, including proofs. An exemption is Chapter IV, which is work early in development (yet valuable to mention). Instead, for Appendix D we provide some considerations about the detection of Continuous Attractors (CANs), which display prominently in Chapters II and III, consideration particularly important in order to avoid confusion when it comes to these concepts, particularly within the experimental neuroscience community.Together, these investigations reveal complementary aspects of how intelligent systems develop useful representations through learning. From biologically plausible learning rules to abstract computational principles, this thesis demonstrates how neural networks can illuminate fundamental mechanisms of intelligence across natural and artificial systems, advancing our understanding of the computational foundations that enable flexible, generalizable learning.
- 일반주제명
- Sparsity
- 일반주제명
- Neurosciences
- 일반주제명
- Neural networks
- 기타저자
- California Institute of Technology Biology and Biological Engineering
- 기본자료저록
- Dissertations Abstracts International. 87-01B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017358779
■00520260202104750
■006m o d
■007cr#unu||||||||
■020 ▼a9798290657219
■035 ▼a(MiAaPQ)AAI32151333
■035 ▼a(MiAaPQ)Caltech17258
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a330
■1001 ▼aVafeidis, Panteleimon.
■24510▼aNeural Network Models of Learning and Generalization
■260 ▼a[Sl]▼bCalifornia Institute of Technology▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a250 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-01, Section: B.
■500 ▼aAdvisor: Rangel, Antonio.
■5021 ▼aThesis (Ph.D.)--California Institute of Technology, 2025.
■520 ▼aNeural networks have emerged as powerful models for understanding both biological and artificial intelligence, yet significant questions remain about how these systems develop rich, generalizable representations of the world. This thesis investigates fundamental principles of learning and generalization across four interconnected domains, bridging insights from theoretical neuroscience and artificial intelligence to advance our understanding of intelligent systems.In Chapter I, we address a central question in associative learning: how do neural circuits learn to associate concepts with one another? We propose a recurrent neural network model incorporating two critical features of cortical architecture-mixed selectivity and compartmentalized neurons. These architectural inductive biases enable a biologically plausible learning rule that achieves stimulus substitution, where neurons respond identically to a conditioned stimulus as they would to the associated unconditioned stimulus. Our model explains a remarkable range of conditioning phenomena under conditions in which traditional associative models fail, highlighting how the cortical architecture may confer significant evolutionary advantages for flexible learning.Chapter II pivots from the static mappings between concepts learned in Chapter I to explore how neural systems develop the precise synaptic connectivity required to establish dynamic mappings for path integration-the ability to maintain an internal sense of direction without external cues. We demonstrate that the same principles of compartmentalized learning can shape networks that accurately track angular position in darkness. Applied to the Drosophilahead direction system, our model develops connectivity patterns strikingly similar to those observed experimentally, with continuous attractor (CAN) dynamics emerging naturally from learning. This offers a novel perspective on how precisely calibrated neural circuits can develop through experience, rather than requiring genetic pre-specification, and explains experimental findings where animals adapt their internal representation when sensory experience changes.In Chapter III, we establish a theoretical framework explaining how disentangled representations-internal models that isolate independent factors of variation in the world-emerge from multi-task learning. We prove that any system competent at multiple related tasks must implicitly represent the underlying latent variables in a linearly decodable form when sufficient tasks are learned. These theoretical guarantees align with experimental results showing neural networks develop generalizable representations when trained on multiple tasks simultaneously. This work reveals a fundamental connection between task diversity and representation quality, with implications for biological cognition and artificial intelligence design, particularly explaining why modern transformer models may develop human-interpretable concepts, and how brains may acquire their impressive zero-shot generalization ability.Chapter IV proposes leveraging Large Language Models as cognitive tools for evaluating latent factor hypotheses for psychology, leveraging the theoretical insights from Chapter III. It suggests that the self-consistency of an LLM's responses given hypothesized psychological factors could serve as a metric for hypothesis evaluation. While preliminary, this approach represents a novel computational methodology for psychology that could transform how hypotheses for human cognition are developed and refined.All chapters are supported by corresponding Appendices that go deeper in particular details, including proofs. An exemption is Chapter IV, which is work early in development (yet valuable to mention). Instead, for Appendix D we provide some considerations about the detection of Continuous Attractors (CANs), which display prominently in Chapters II and III, consideration particularly important in order to avoid confusion when it comes to these concepts, particularly within the experimental neuroscience community.Together, these investigations reveal complementary aspects of how intelligent systems develop useful representations through learning. From biologically plausible learning rules to abstract computational principles, this thesis demonstrates how neural networks can illuminate fundamental mechanisms of intelligence across natural and artificial systems, advancing our understanding of the computational foundations that enable flexible, generalizable learning.
■590 ▼aSchool code: 0037.
■650 4▼aSparsity
■650 4▼aNeurosciences
■650 4▼aNeural networks
■690 ▼a0800
■690 ▼a0317
■71020▼aCalifornia Institute of Technology▼bBiology and Biological Engineering.
■7730 ▼tDissertations Abstracts International▼g87-01B.
■790 ▼a0037
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17358779▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


