서브메뉴
검색
Learning in Large Neural Networks
Learning in Large Neural Networks
Detailed Information
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103547
- ISBN
- 9798280710436
- DDC
- 519
- 저자명
- Bordelon, Blake.
- 서명/저자
- Learning in Large Neural Networks
- 발행사항
- [Sl] : Harvard University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 1396 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-12, Section: B.
- 주기사항
- Advisor: Pehlevan, Cengiz.
- 학위논문주기
- Thesis (Ph.D.)--Harvard University, 2025.
- 초록/해제
- 요약In this thesis, I will summarize my recent works on theoretical frameworks for learning, generalization, scaling limits, and scaling laws of large neural networks. In the first part of the thesis, we will examine the kinds of limits attained by training randomly initialized networks. The two limits of primary interest are the infinite width and infinite depth feature-learning limits.The large width limit will take the form of a dynamical mean field theory (DMFT), where neurons asymptotically decouple and the macroscopic dynamics of the network are governed by population averages over the neurons in each hidden layer. This theory computes the dynamics of the learned representations of data in each hidden layer of the network throughout training as well as the dynamics of the network output predictions. Asymptotic corrections to this mean field limit will be computed from fluctuations around the DMFT saddle point. The mean field limit will then be stressed tested on several realistic networks such as convolutional networks and transformers on realistic computer vision and language modeling datasets. We next investigate the infinite depth limit of residual neural networks, where each hidden layer is a trainable perturbation of the identity map. When the residual branches are scaled correctly, these models admit infinite width and depth limits which are computable with a DMFT where intermediate layers approach a continuum limit described by the solution to a set of stochastic integral equations. We empirically show that scaling the depth and width correctly enable hyperparameter transfer, where optimal hyperparameters (learning rates, batch sizes, momentum values, etc) in small width and small depth models are the same in large width and large depth models, reducing the need for hyperparamter tuning. We will extend these mathematical techniques to analyze several distinct infinite-parameter limits for transformer models.Next, I will describe simplified models of learning which enable a statistical analysis of generalization in a data-limited or parameter-limited regime. As in the previous section, we present both static and dynamic versions of these results. For data distributions and architectures that generate power law spectra for their limiting kernels, these theories provide predictions for the scaling laws with respect to the key computational and statistical resources: training time, model parameters, and total available data. I show that this setting provides a toy model of compute optimal scaling laws where model size and training time are traded off optimally. Lastly, we use similar mathematical techniques to analyze a toy model of hyperparameter transfer in randomly initialized deep linear networks.Lastly, we move beyond randomly initialized deep networks and attempt to address how structured neural representations, such as the cortical representations of external stimuli in real brains, encode an implicit learning bias. We start with a simple neural circuit trained with the delta rule, finding that the spectral decomposition of the population code controls which learning tasks can be learned in a sample-efficient manner. Next, we examine how different learning rules alter feature learning dynamics and inductive bias in multilayer neural networks, extending the DMFT for multilayer networks to other biologically plausible learning rules. Lastly, we illustrate how the geometry of neural codes which respect symmetries in the data generating process controls the manifold capacity of linear readouts.
- 일반주제명
- Applied mathematics
- 일반주제명
- Computer science
- 일반주제명
- Statistical physics
- 키워드
- Deep learning
- 키워드
- Scaling laws
- 기타저자
- Harvard University Engineering and Applied Sciences - Applied Math
- 기본자료저록
- Dissertations Abstracts International. 86-12B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017357692
■00520260202103547
■006m o d
■007cr#unu||||||||
■020 ▼a9798280710436
■035 ▼a(MiAaPQ)AAI32041386
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a519
■1001 ▼aBordelon, Blake.▼0(orcid)0000-0003-0455-9445
■24510▼aLearning in Large Neural Networks
■260 ▼a[Sl]▼bHarvard University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a1396 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-12, Section: B.
■500 ▼aAdvisor: Pehlevan, Cengiz.
■5021 ▼aThesis (Ph.D.)--Harvard University, 2025.
■520 ▼aIn this thesis, I will summarize my recent works on theoretical frameworks for learning, generalization, scaling limits, and scaling laws of large neural networks. In the first part of the thesis, we will examine the kinds of limits attained by training randomly initialized networks. The two limits of primary interest are the infinite width and infinite depth feature-learning limits.The large width limit will take the form of a dynamical mean field theory (DMFT), where neurons asymptotically decouple and the macroscopic dynamics of the network are governed by population averages over the neurons in each hidden layer. This theory computes the dynamics of the learned representations of data in each hidden layer of the network throughout training as well as the dynamics of the network output predictions. Asymptotic corrections to this mean field limit will be computed from fluctuations around the DMFT saddle point. The mean field limit will then be stressed tested on several realistic networks such as convolutional networks and transformers on realistic computer vision and language modeling datasets. We next investigate the infinite depth limit of residual neural networks, where each hidden layer is a trainable perturbation of the identity map. When the residual branches are scaled correctly, these models admit infinite width and depth limits which are computable with a DMFT where intermediate layers approach a continuum limit described by the solution to a set of stochastic integral equations. We empirically show that scaling the depth and width correctly enable hyperparameter transfer, where optimal hyperparameters (learning rates, batch sizes, momentum values, etc) in small width and small depth models are the same in large width and large depth models, reducing the need for hyperparamter tuning. We will extend these mathematical techniques to analyze several distinct infinite-parameter limits for transformer models.Next, I will describe simplified models of learning which enable a statistical analysis of generalization in a data-limited or parameter-limited regime. As in the previous section, we present both static and dynamic versions of these results. For data distributions and architectures that generate power law spectra for their limiting kernels, these theories provide predictions for the scaling laws with respect to the key computational and statistical resources: training time, model parameters, and total available data. I show that this setting provides a toy model of compute optimal scaling laws where model size and training time are traded off optimally. Lastly, we use similar mathematical techniques to analyze a toy model of hyperparameter transfer in randomly initialized deep linear networks.Lastly, we move beyond randomly initialized deep networks and attempt to address how structured neural representations, such as the cortical representations of external stimuli in real brains, encode an implicit learning bias. We start with a simple neural circuit trained with the delta rule, finding that the spectral decomposition of the population code controls which learning tasks can be learned in a sample-efficient manner. Next, we examine how different learning rules alter feature learning dynamics and inductive bias in multilayer neural networks, extending the DMFT for multilayer networks to other biologically plausible learning rules. Lastly, we illustrate how the geometry of neural codes which respect symmetries in the data generating process controls the manifold capacity of linear readouts.
■590 ▼aSchool code: 0084.
■650 4▼aApplied mathematics
■650 4▼aComputer science
■650 4▼aStatistical physics
■653 ▼aDeep learning
■653 ▼aScaling laws
■653 ▼aLarge neural networks
■653 ▼aConvolutional networks
■690 ▼a0364
■690 ▼a0984
■690 ▼a0217
■71020▼aHarvard University▼bEngineering and Applied Sciences - Applied Math.
■7730 ▼tDissertations Abstracts International▼g86-12B.
■790 ▼a0084
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357692▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.
Preview
Export
ChatGPT Discussion
AI Recommended Related Books
Подробнее информация.
- Бронирование
- не существует
- моя папка
- Первый запрос зрения
- Non-Book Loan Application
- Nighttime Book Loan Application
Available after logging in.


