본문

서브메뉴

Learning in Large Neural Networks
Learning in Large Neural Networks
Learning in Large Neural Networks

Detailed Information

자료유형  
 학위논문 서양
최종처리일시  
20260202103547
ISBN  
9798280710436
DDC  
519
저자명  
Bordelon, Blake.
서명/저자  
Learning in Large Neural Networks
발행사항  
[Sl] : Harvard University, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
1396 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-12, Section: B.
주기사항  
Advisor: Pehlevan, Cengiz.
학위논문주기  
Thesis (Ph.D.)--Harvard University, 2025.
초록/해제  
요약In this thesis, I will summarize my recent works on theoretical frameworks for learning, generalization, scaling limits, and scaling laws of large neural networks. In the first part of the thesis, we will examine the kinds of limits attained by training randomly initialized networks. The two limits of primary interest are the infinite width and infinite depth feature-learning limits.The large width limit will take the form of a dynamical mean field theory (DMFT), where neurons asymptotically decouple and the macroscopic dynamics of the network are governed by population averages over the neurons in each hidden layer. This theory computes the dynamics of the learned representations of data in each hidden layer of the network throughout training as well as the dynamics of the network output predictions. Asymptotic corrections to this mean field limit will be computed from fluctuations around the DMFT saddle point. The mean field limit will then be stressed tested on several realistic networks such as convolutional networks and transformers on realistic computer vision and language modeling datasets. We next investigate the infinite depth limit of residual neural networks, where each hidden layer is a trainable perturbation of the identity map. When the residual branches are scaled correctly, these models admit infinite width and depth limits which are computable with a DMFT where intermediate layers approach a continuum limit described by the solution to a set of stochastic integral equations. We empirically show that scaling the depth and width correctly enable hyperparameter transfer, where optimal hyperparameters (learning rates, batch sizes, momentum values, etc) in small width and small depth models are the same in large width and large depth models, reducing the need for hyperparamter tuning. We will extend these mathematical techniques to analyze several distinct infinite-parameter limits for transformer models.Next, I will describe simplified models of learning which enable a statistical analysis of generalization in a data-limited or parameter-limited regime. As in the previous section, we present both static and dynamic versions of these results. For data distributions and architectures that generate power law spectra for their limiting kernels, these theories provide predictions for the scaling laws with respect to the key computational and statistical resources: training time, model parameters, and total available data. I show that this setting provides a toy model of compute optimal scaling laws where model size and training time are traded off optimally. Lastly, we use similar mathematical techniques to analyze a toy model of hyperparameter transfer in randomly initialized deep linear networks.Lastly, we move beyond randomly initialized deep networks and attempt to address how structured neural representations, such as the cortical representations of external stimuli in real brains, encode an implicit learning bias. We start with a simple neural circuit trained with the delta rule, finding that the spectral decomposition of the population code controls which learning tasks can be learned in a sample-efficient manner. Next, we examine how different learning rules alter feature learning dynamics and inductive bias in multilayer neural networks, extending the DMFT for multilayer networks to other biologically plausible learning rules. Lastly, we illustrate how the geometry of neural codes which respect symmetries in the data generating process controls the manifold capacity of linear readouts.
일반주제명  
Applied mathematics
일반주제명  
Computer science
일반주제명  
Statistical physics
키워드  
Deep learning
키워드  
Scaling laws
키워드  
Large neural networks
키워드  
Convolutional networks
기타저자  
Harvard University Engineering and Applied Sciences - Applied Math
기본자료저록  
Dissertations Abstracts International. 86-12B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017357692
■00520260202103547
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798280710436
■035    ▼a(MiAaPQ)AAI32041386
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a519
■1001  ▼aBordelon,  Blake.▼0(orcid)0000-0003-0455-9445
■24510▼aLearning  in  Large  Neural  Networks
■260    ▼a[Sl]▼bHarvard  University▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a1396  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-12,  Section:  B.
■500    ▼aAdvisor:  Pehlevan,  Cengiz.
■5021  ▼aThesis  (Ph.D.)--Harvard  University,  2025.
■520    ▼aIn  this  thesis,  I  will  summarize  my  recent  works  on  theoretical  frameworks  for  learning,  generalization,  scaling  limits,  and  scaling  laws  of  large  neural  networks.  In  the  first  part  of  the  thesis,  we  will  examine  the  kinds  of  limits  attained  by  training  randomly  initialized  networks.  The  two  limits  of  primary  interest  are  the  infinite  width  and  infinite  depth  feature-learning  limits.The  large  width  limit  will  take  the  form  of  a  dynamical  mean  field  theory  (DMFT),  where  neurons  asymptotically  decouple  and  the  macroscopic  dynamics  of  the  network  are  governed  by  population  averages  over  the  neurons  in  each  hidden  layer.  This  theory  computes  the  dynamics  of  the  learned  representations  of  data  in  each  hidden  layer  of  the  network  throughout  training  as  well  as  the  dynamics  of  the  network  output  predictions.  Asymptotic  corrections  to  this  mean  field  limit  will  be  computed  from  fluctuations  around  the  DMFT  saddle  point.  The  mean  field  limit  will  then  be  stressed  tested  on  several  realistic  networks  such  as  convolutional  networks  and  transformers  on  realistic  computer  vision  and  language  modeling  datasets.  We  next  investigate  the  infinite  depth  limit  of  residual  neural  networks,  where  each  hidden  layer  is  a  trainable  perturbation  of  the  identity  map.  When  the  residual  branches  are  scaled  correctly,  these  models  admit  infinite  width  and  depth  limits  which  are  computable  with  a  DMFT  where  intermediate  layers  approach  a  continuum  limit  described  by  the  solution  to  a  set  of  stochastic  integral  equations.  We  empirically  show  that  scaling  the  depth  and  width  correctly  enable  hyperparameter  transfer,  where  optimal  hyperparameters  (learning  rates,  batch  sizes,  momentum  values,  etc)  in  small  width  and  small  depth  models  are  the  same  in  large  width  and  large  depth  models,  reducing  the  need  for  hyperparamter  tuning.  We  will  extend  these  mathematical  techniques  to  analyze  several  distinct  infinite-parameter  limits  for  transformer  models.Next,  I  will  describe  simplified  models  of  learning  which  enable  a  statistical  analysis  of  generalization  in  a  data-limited  or  parameter-limited  regime.  As  in  the  previous  section,  we  present  both  static  and  dynamic  versions  of  these  results.  For  data  distributions  and  architectures  that  generate  power  law  spectra  for  their  limiting  kernels,  these  theories  provide  predictions  for  the  scaling  laws  with  respect  to  the  key  computational  and  statistical  resources:  training  time,  model  parameters,  and  total  available  data.  I  show  that  this  setting  provides  a  toy  model  of  compute  optimal  scaling  laws  where  model  size  and  training  time  are  traded  off  optimally.  Lastly,  we  use  similar  mathematical  techniques  to  analyze  a  toy  model  of  hyperparameter  transfer  in  randomly  initialized  deep  linear  networks.Lastly,  we  move  beyond  randomly  initialized  deep  networks  and  attempt  to  address  how  structured  neural  representations,  such  as  the  cortical  representations  of  external  stimuli  in  real  brains,  encode  an  implicit  learning  bias.  We  start  with  a  simple  neural  circuit  trained  with  the  delta  rule,  finding  that  the  spectral  decomposition  of  the  population  code  controls  which  learning  tasks  can  be  learned  in  a  sample-efficient  manner.  Next,  we  examine  how  different  learning  rules  alter  feature  learning  dynamics  and  inductive  bias  in  multilayer  neural  networks,  extending  the  DMFT  for  multilayer  networks  to  other  biologically  plausible  learning  rules.  Lastly,  we  illustrate  how  the  geometry  of  neural  codes  which  respect  symmetries  in  the  data  generating  process  controls  the  manifold  capacity  of  linear  readouts.
■590    ▼aSchool  code:  0084.
■650  4▼aApplied  mathematics
■650  4▼aComputer  science
■650  4▼aStatistical  physics
■653    ▼aDeep  learning
■653    ▼aScaling  laws
■653    ▼aLarge  neural  networks
■653    ▼aConvolutional  networks
■690    ▼a0364
■690    ▼a0984
■690    ▼a0217
■71020▼aHarvard  University▼bEngineering  and  Applied  Sciences  -  Applied  Math.
■7730  ▼tDissertations  Abstracts  International▼g86-12B.
■790    ▼a0084
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357692▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

Preview

Export

ChatGPT Discussion

AI Recommended Related Books


    New Books MORE
    Statistics for the past 3 years. Go to brief

    Подробнее информация.

    • Бронирование
    • не существует
    • моя папка
    • Первый запрос зрения
    • Non-Book Loan Application
    • Nighttime Book Loan Application
    материал
    Reg No. Количество платежных Местоположение статус Ленд информации
    TF18841 전자도서 대출가능 My Folder 부재도서신고 비도서대출신청 야간 도서대출신청

    * Бронирование доступны в заимствований книги. Чтобы сделать предварительный заказ, пожалуйста, нажмите кнопку бронирование

    Books borrowed together with this book

    Related Popular Books

    Available after logging in.