본문

서브메뉴

Learning Mechanics of Neural Networks: Conservation Laws, Implicit Bias, and Feature Learning
Learning Mechanics of Neural Networks: Conservation Laws, Implicit Bias, and Feature Learn...
Learning Mechanics of Neural Networks: Conservation Laws, Implicit Bias, and Feature Learning

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202104850
ISBN  
9798288814761
DDC  
000
저자명  
Kunin, Daniel.
서명/저자  
Learning Mechanics of Neural Networks: Conservation Laws, Implicit Bias, and Feature Learning
발행사항  
[Sl] : Stanford University, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
185 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-02, Section: B.
주기사항  
Advisor: Ganguli, Surya.
학위논문주기  
Thesis (Ph.D.)--Stanford University, 2025.
초록/해제  
요약Deep learning has revolutionized artificial intelligence (AI), achieving superhuman performance in tasks from visual recognition to natural language processing, largely due to increasing computational scale. While deeper models, larger datasets, and longer training will undoubtedly continue to yield performance gains, a fundamental question remains: what is the mathematical basis for the practical success of AI? Today's AI systems learn through a process that remains largely mysterious to us, creating both opportunities and serious challenges as these technologies grow more powerful and more deeply embedded in society. To overcome these challenges, we must understand the core mathematical principles that underpin learning. This understanding not only promises significant advancements in AI, but also has the potential to drive major scientific breakthroughs in uncovering the structure of natural intelligence. AI and neuroscience have long shared a history of crosspollination, with progress in one field often driving innovation in the other. Throughout my Ph.D., I have drawn on ideas from statistics, physics, and neuroscience to uncover the mathematical principles of learning in artificial and natural intelligence. This thesis presents the results of that effort.Two prevailing perspectives. To understand learning, we must unravel an intricate interaction between a network's architecture, a training dataset, and an optimization strategy. My research approaches this interaction from two complementary perspectives:• The feature learning perspective studies how the structure and parameterization of a network enable it to extract and compose task-relevant features from data. • The implicit bias perspective examines how hyperparameters of the optimization process, such as the learning rate and batch size, implicitly guide the network toward simple solutions.This thesis is organized into two chapters, each aligned with one of the two perspectives. Each chapter synthesizes results from three research papers, with key derivations included for clarity in a shared appendix. Technical details and extended results are deferred to the original publications.Feature learning perspective. The impressive performance of neural networks has been attributed to their ability to extract task-relevant representation from data, a process termed feature learning. Notably, AI systems often learn representations similar to those seen in biological systems. However, the mechanisms underlying feature learning remain largely unknown. In this chapter, we investigate the emergence of feature learning through the following three studies:1. We derive exact solutions to a minimal model that transitions between lazy and rich learning, precisely elucidating how unbalanced initialization variances and learning rates determine the degree of feature learning in a finite-width network. This work, co-first authored with Allan Raventos, appeared at NeurIPS 2024 Kunin et al. [2024]. 2. We introduce Alternating Gradient Flows, a framework modeling feature learning in two-layer networks with small initialization as utility maximization and cost minimization-unifying saddle-to-saddle analyses and explaining the emergence of Fourier features. This first authored work is currently under review Kunin et al. [2025]. 3. We identify a late-stage tradeoff between margin maximization and asymmetric norm minimization that promotes feature learning in networks with homogeneous activations-potentially degrading robustness and explaining Neural Collapse. This work, co-first authored with Atsushi Yamamura, was published at ICLR 2023 Kunin et al. [2023b].Implicit bias perspective. Contrary to traditional statistical learning theory, neural networks can generalize remarkably well despite being trained past the point at which they interpolate their training data. This unexpected behavior suggests the existence of inductive biases that regularize the network to find low-complexity solutions when available. In this chapter, we take the following steps to investigate the role of implicit bias in deep learning:1. We exploit architectural symmetry to analytically describe the learning dynamics of various parameter combinations at finite learning rates and batch sizes. This work, co-first authored with Hidenori Tanaka, was published at ICLR 2021 Kunin et al. [2021]. 2. We use tools from statistical physics to identify how noisy gradients lead to oscillatory behavior in the limiting dynamics of neural networks leading to anomalous diffusion in parameter space. This work, co-first authored with Javier Sagastuy-Brena, was published in Neural Computation Kunin et al. [2023a] and included with permission from MIT Press. 3. We reveal how stochasticity from mini-batch gradients biases overparameterized neural networks towards "invariant sets" corresponding to simpler subnetworks with improved generalization. This work, co-first authored with Feng Chen and Atsushi Yamamura, was published in the Journal of Statistical Mechanics Chen et al. [2024].Together, the results presented in this thesis advance our mathematical understanding of learning in neural networks, highlighting how data, architecture, and optimization interact to shape generalization.
일반주제명  
Conservation laws
일반주제명  
Deep learning
일반주제명  
Coordinate transformations
일반주제명  
Geometry
일반주제명  
Neural networks
일반주제명  
Symmetry
일반주제명  
Cognitive psychology
키워드  
Neural networks
키워드  
Deep learning
기타저자  
Stanford University.
기본자료저록  
Dissertations Abstracts International. 87-02B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017359211
■00520260202104850
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798288814761
■035    ▼a(MiAaPQ)AAI32200948
■035    ▼a(MiAaPQ)Stanfordgs143yc6699
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a000
■1001  ▼aKunin,  Daniel.
■24510▼aLearning  Mechanics  of  Neural  Networks:  Conservation  Laws,  Implicit  Bias,  and  Feature  Learning
■260    ▼a[Sl]▼bStanford  University▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a185  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-02,  Section:  B.
■500    ▼aAdvisor:  Ganguli,  Surya.
■5021  ▼aThesis  (Ph.D.)--Stanford  University,  2025.
■520    ▼aDeep  learning  has  revolutionized  artificial  intelligence  (AI),  achieving  superhuman  performance  in  tasks  from  visual  recognition  to  natural  language  processing,  largely  due  to  increasing  computational  scale.  While  deeper  models,  larger  datasets,  and  longer  training  will  undoubtedly  continue  to  yield  performance  gains,  a  fundamental  question  remains:  what  is  the  mathematical  basis  for  the  practical  success  of  AI?  Today's  AI  systems  learn  through  a  process  that  remains  largely  mysterious  to  us,  creating  both  opportunities  and  serious  challenges  as  these  technologies  grow  more  powerful  and  more  deeply  embedded  in  society.  To  overcome  these  challenges,  we  must  understand  the  core  mathematical  principles  that  underpin  learning.  This  understanding  not  only  promises  significant  advancements  in  AI,  but  also  has  the  potential  to  drive  major  scientific  breakthroughs  in  uncovering  the  structure  of  natural  intelligence.  AI  and  neuroscience  have  long  shared  a  history  of  crosspollination,  with  progress  in  one  field  often  driving  innovation  in  the  other.  Throughout  my  Ph.D.,  I  have  drawn  on  ideas  from  statistics,  physics,  and  neuroscience  to  uncover  the  mathematical  principles  of  learning  in  artificial  and  natural  intelligence.  This  thesis  presents  the  results  of  that  effort.Two  prevailing  perspectives.  To  understand  learning,  we  must  unravel  an  intricate  interaction  between  a  network's  architecture,  a  training  dataset,  and  an  optimization  strategy.  My  research  approaches  this  interaction  from  two  complementary  perspectives:•  The  feature  learning  perspective  studies  how  the  structure  and  parameterization  of  a  network  enable  it  to  extract  and  compose  task-relevant  features  from  data.  •  The  implicit  bias  perspective  examines  how  hyperparameters  of  the  optimization  process,  such  as  the  learning  rate  and  batch  size,  implicitly  guide  the  network  toward  simple  solutions.This  thesis  is  organized  into  two  chapters,  each  aligned  with  one  of  the  two  perspectives.  Each  chapter  synthesizes  results  from  three  research  papers,  with  key  derivations  included  for  clarity  in  a  shared  appendix.  Technical  details  and  extended  results  are  deferred  to  the  original  publications.Feature  learning  perspective.  The  impressive  performance  of  neural  networks  has  been  attributed  to  their  ability  to  extract  task-relevant  representation  from  data,  a  process  termed  feature  learning.  Notably,  AI  systems  often  learn  representations  similar  to  those  seen  in  biological  systems.  However,  the  mechanisms  underlying  feature  learning  remain  largely  unknown.  In  this  chapter,  we  investigate  the  emergence  of  feature  learning  through  the  following  three  studies:1.  We  derive  exact  solutions  to  a  minimal  model  that  transitions  between  lazy  and  rich  learning,  precisely  elucidating  how  unbalanced  initialization  variances  and  learning  rates  determine  the  degree  of  feature  learning  in  a  finite-width  network.  This  work,  co-first  authored  with  Allan  Raventos,  appeared  at  NeurIPS  2024  Kunin  et  al.  [2024].  2.  We  introduce  Alternating  Gradient  Flows,  a  framework  modeling  feature  learning  in  two-layer  networks  with  small  initialization  as  utility  maximization  and  cost  minimization-unifying  saddle-to-saddle  analyses  and  explaining  the  emergence  of  Fourier  features.  This  first  authored  work  is  currently  under  review  Kunin  et  al.  [2025].  3.  We  identify  a  late-stage  tradeoff  between  margin  maximization  and  asymmetric  norm  minimization  that  promotes  feature  learning  in  networks  with  homogeneous  activations-potentially  degrading  robustness  and  explaining  Neural  Collapse.  This  work,  co-first  authored  with  Atsushi  Yamamura,  was  published  at  ICLR  2023  Kunin  et  al.  [2023b].Implicit  bias  perspective.  Contrary  to  traditional  statistical  learning  theory,  neural  networks  can  generalize  remarkably  well  despite  being  trained  past  the  point  at  which  they  interpolate  their  training  data.  This  unexpected  behavior  suggests  the  existence  of  inductive  biases  that  regularize  the  network  to  find  low-complexity  solutions  when  available.  In  this  chapter,  we  take  the  following  steps  to  investigate  the  role  of  implicit  bias  in  deep  learning:1.  We  exploit  architectural  symmetry  to  analytically  describe  the  learning  dynamics  of  various  parameter  combinations  at  finite  learning  rates  and  batch  sizes.  This  work,  co-first  authored  with  Hidenori  Tanaka,  was  published  at  ICLR  2021  Kunin  et  al.  [2021].  2.  We  use  tools  from  statistical  physics  to  identify  how  noisy  gradients  lead  to  oscillatory  behavior  in  the  limiting  dynamics  of  neural  networks  leading  to  anomalous  diffusion  in  parameter  space.  This  work,  co-first  authored  with  Javier  Sagastuy-Brena,  was  published  in  Neural  Computation  Kunin  et  al.  [2023a]  and  included  with  permission  from  MIT  Press.  3.  We  reveal  how  stochasticity  from  mini-batch  gradients  biases  overparameterized  neural  networks  towards  "invariant  sets"  corresponding  to  simpler  subnetworks  with  improved  generalization.  This  work,  co-first  authored  with  Feng  Chen  and  Atsushi  Yamamura,  was  published  in  the  Journal  of  Statistical  Mechanics  Chen  et  al.  [2024].Together,  the  results  presented  in  this  thesis  advance  our  mathematical  understanding  of  learning  in  neural  networks,  highlighting  how  data,  architecture,  and  optimization  interact  to  shape  generalization.
■590    ▼aSchool  code:  0212.
■650  4▼aConservation  laws
■650  4▼aDeep  learning
■650  4▼aCoordinate  transformations
■650  4▼aGeometry
■650  4▼aNeural  networks
■650  4▼aSymmetry
■650  4▼aCognitive  psychology
■653    ▼aNeural  networks
■653    ▼aDeep  learning
■690    ▼a0800
■690    ▼a0633
■71020▼aStanford  University.
■7730  ▼tDissertations  Abstracts  International▼g87-02B.
■790    ▼a0212
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17359211▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF14738 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.