서브메뉴
검색
Learning Mechanics of Neural Networks: Conservation Laws, Implicit Bias, and Feature Learning
Learning Mechanics of Neural Networks: Conservation Laws, Implicit Bias, and Feature Learning
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202104850
- ISBN
- 9798288814761
- DDC
- 000
- 저자명
- Kunin, Daniel.
- 서명/저자
- Learning Mechanics of Neural Networks: Conservation Laws, Implicit Bias, and Feature Learning
- 발행사항
- [Sl] : Stanford University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 185 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-02, Section: B.
- 주기사항
- Advisor: Ganguli, Surya.
- 학위논문주기
- Thesis (Ph.D.)--Stanford University, 2025.
- 초록/해제
- 요약Deep learning has revolutionized artificial intelligence (AI), achieving superhuman performance in tasks from visual recognition to natural language processing, largely due to increasing computational scale. While deeper models, larger datasets, and longer training will undoubtedly continue to yield performance gains, a fundamental question remains: what is the mathematical basis for the practical success of AI? Today's AI systems learn through a process that remains largely mysterious to us, creating both opportunities and serious challenges as these technologies grow more powerful and more deeply embedded in society. To overcome these challenges, we must understand the core mathematical principles that underpin learning. This understanding not only promises significant advancements in AI, but also has the potential to drive major scientific breakthroughs in uncovering the structure of natural intelligence. AI and neuroscience have long shared a history of crosspollination, with progress in one field often driving innovation in the other. Throughout my Ph.D., I have drawn on ideas from statistics, physics, and neuroscience to uncover the mathematical principles of learning in artificial and natural intelligence. This thesis presents the results of that effort.Two prevailing perspectives. To understand learning, we must unravel an intricate interaction between a network's architecture, a training dataset, and an optimization strategy. My research approaches this interaction from two complementary perspectives:• The feature learning perspective studies how the structure and parameterization of a network enable it to extract and compose task-relevant features from data. • The implicit bias perspective examines how hyperparameters of the optimization process, such as the learning rate and batch size, implicitly guide the network toward simple solutions.This thesis is organized into two chapters, each aligned with one of the two perspectives. Each chapter synthesizes results from three research papers, with key derivations included for clarity in a shared appendix. Technical details and extended results are deferred to the original publications.Feature learning perspective. The impressive performance of neural networks has been attributed to their ability to extract task-relevant representation from data, a process termed feature learning. Notably, AI systems often learn representations similar to those seen in biological systems. However, the mechanisms underlying feature learning remain largely unknown. In this chapter, we investigate the emergence of feature learning through the following three studies:1. We derive exact solutions to a minimal model that transitions between lazy and rich learning, precisely elucidating how unbalanced initialization variances and learning rates determine the degree of feature learning in a finite-width network. This work, co-first authored with Allan Raventos, appeared at NeurIPS 2024 Kunin et al. [2024]. 2. We introduce Alternating Gradient Flows, a framework modeling feature learning in two-layer networks with small initialization as utility maximization and cost minimization-unifying saddle-to-saddle analyses and explaining the emergence of Fourier features. This first authored work is currently under review Kunin et al. [2025]. 3. We identify a late-stage tradeoff between margin maximization and asymmetric norm minimization that promotes feature learning in networks with homogeneous activations-potentially degrading robustness and explaining Neural Collapse. This work, co-first authored with Atsushi Yamamura, was published at ICLR 2023 Kunin et al. [2023b].Implicit bias perspective. Contrary to traditional statistical learning theory, neural networks can generalize remarkably well despite being trained past the point at which they interpolate their training data. This unexpected behavior suggests the existence of inductive biases that regularize the network to find low-complexity solutions when available. In this chapter, we take the following steps to investigate the role of implicit bias in deep learning:1. We exploit architectural symmetry to analytically describe the learning dynamics of various parameter combinations at finite learning rates and batch sizes. This work, co-first authored with Hidenori Tanaka, was published at ICLR 2021 Kunin et al. [2021]. 2. We use tools from statistical physics to identify how noisy gradients lead to oscillatory behavior in the limiting dynamics of neural networks leading to anomalous diffusion in parameter space. This work, co-first authored with Javier Sagastuy-Brena, was published in Neural Computation Kunin et al. [2023a] and included with permission from MIT Press. 3. We reveal how stochasticity from mini-batch gradients biases overparameterized neural networks towards "invariant sets" corresponding to simpler subnetworks with improved generalization. This work, co-first authored with Feng Chen and Atsushi Yamamura, was published in the Journal of Statistical Mechanics Chen et al. [2024].Together, the results presented in this thesis advance our mathematical understanding of learning in neural networks, highlighting how data, architecture, and optimization interact to shape generalization.
- 일반주제명
- Conservation laws
- 일반주제명
- Deep learning
- 일반주제명
- Geometry
- 일반주제명
- Neural networks
- 일반주제명
- Symmetry
- 일반주제명
- Cognitive psychology
- 키워드
- Neural networks
- 키워드
- Deep learning
- 기타저자
- Stanford University.
- 기본자료저록
- Dissertations Abstracts International. 87-02B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017359211
■00520260202104850
■006m o d
■007cr#unu||||||||
■020 ▼a9798288814761
■035 ▼a(MiAaPQ)AAI32200948
■035 ▼a(MiAaPQ)Stanfordgs143yc6699
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a000
■1001 ▼aKunin, Daniel.
■24510▼aLearning Mechanics of Neural Networks: Conservation Laws, Implicit Bias, and Feature Learning
■260 ▼a[Sl]▼bStanford University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a185 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-02, Section: B.
■500 ▼aAdvisor: Ganguli, Surya.
■5021 ▼aThesis (Ph.D.)--Stanford University, 2025.
■520 ▼aDeep learning has revolutionized artificial intelligence (AI), achieving superhuman performance in tasks from visual recognition to natural language processing, largely due to increasing computational scale. While deeper models, larger datasets, and longer training will undoubtedly continue to yield performance gains, a fundamental question remains: what is the mathematical basis for the practical success of AI? Today's AI systems learn through a process that remains largely mysterious to us, creating both opportunities and serious challenges as these technologies grow more powerful and more deeply embedded in society. To overcome these challenges, we must understand the core mathematical principles that underpin learning. This understanding not only promises significant advancements in AI, but also has the potential to drive major scientific breakthroughs in uncovering the structure of natural intelligence. AI and neuroscience have long shared a history of crosspollination, with progress in one field often driving innovation in the other. Throughout my Ph.D., I have drawn on ideas from statistics, physics, and neuroscience to uncover the mathematical principles of learning in artificial and natural intelligence. This thesis presents the results of that effort.Two prevailing perspectives. To understand learning, we must unravel an intricate interaction between a network's architecture, a training dataset, and an optimization strategy. My research approaches this interaction from two complementary perspectives:• The feature learning perspective studies how the structure and parameterization of a network enable it to extract and compose task-relevant features from data. • The implicit bias perspective examines how hyperparameters of the optimization process, such as the learning rate and batch size, implicitly guide the network toward simple solutions.This thesis is organized into two chapters, each aligned with one of the two perspectives. Each chapter synthesizes results from three research papers, with key derivations included for clarity in a shared appendix. Technical details and extended results are deferred to the original publications.Feature learning perspective. The impressive performance of neural networks has been attributed to their ability to extract task-relevant representation from data, a process termed feature learning. Notably, AI systems often learn representations similar to those seen in biological systems. However, the mechanisms underlying feature learning remain largely unknown. In this chapter, we investigate the emergence of feature learning through the following three studies:1. We derive exact solutions to a minimal model that transitions between lazy and rich learning, precisely elucidating how unbalanced initialization variances and learning rates determine the degree of feature learning in a finite-width network. This work, co-first authored with Allan Raventos, appeared at NeurIPS 2024 Kunin et al. [2024]. 2. We introduce Alternating Gradient Flows, a framework modeling feature learning in two-layer networks with small initialization as utility maximization and cost minimization-unifying saddle-to-saddle analyses and explaining the emergence of Fourier features. This first authored work is currently under review Kunin et al. [2025]. 3. We identify a late-stage tradeoff between margin maximization and asymmetric norm minimization that promotes feature learning in networks with homogeneous activations-potentially degrading robustness and explaining Neural Collapse. This work, co-first authored with Atsushi Yamamura, was published at ICLR 2023 Kunin et al. [2023b].Implicit bias perspective. Contrary to traditional statistical learning theory, neural networks can generalize remarkably well despite being trained past the point at which they interpolate their training data. This unexpected behavior suggests the existence of inductive biases that regularize the network to find low-complexity solutions when available. In this chapter, we take the following steps to investigate the role of implicit bias in deep learning:1. We exploit architectural symmetry to analytically describe the learning dynamics of various parameter combinations at finite learning rates and batch sizes. This work, co-first authored with Hidenori Tanaka, was published at ICLR 2021 Kunin et al. [2021]. 2. We use tools from statistical physics to identify how noisy gradients lead to oscillatory behavior in the limiting dynamics of neural networks leading to anomalous diffusion in parameter space. This work, co-first authored with Javier Sagastuy-Brena, was published in Neural Computation Kunin et al. [2023a] and included with permission from MIT Press. 3. We reveal how stochasticity from mini-batch gradients biases overparameterized neural networks towards "invariant sets" corresponding to simpler subnetworks with improved generalization. This work, co-first authored with Feng Chen and Atsushi Yamamura, was published in the Journal of Statistical Mechanics Chen et al. [2024].Together, the results presented in this thesis advance our mathematical understanding of learning in neural networks, highlighting how data, architecture, and optimization interact to shape generalization.
■590 ▼aSchool code: 0212.
■650 4▼aConservation laws
■650 4▼aDeep learning
■650 4▼aCoordinate transformations
■650 4▼aGeometry
■650 4▼aNeural networks
■650 4▼aSymmetry
■650 4▼aCognitive psychology
■653 ▼aNeural networks
■653 ▼aDeep learning
■690 ▼a0800
■690 ▼a0633
■71020▼aStanford University.
■7730 ▼tDissertations Abstracts International▼g87-02B.
■790 ▼a0212
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17359211▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


