서브메뉴
검색
Scaling and Renormalization in Statistical Learning
Scaling and Renormalization in Statistical Learning
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211152830
- ISBN
- 9798346571490
- DDC
- 530.1
- 서명/저자
- Scaling and Renormalization in Statistical Learning
- 발행사항
- [Sl] : Harvard University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 488 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-05, Section: B.
- 주기사항
- Advisor: Pehlevan, Cengiz.
- 학위논문주기
- Thesis (Ph.D.)--Harvard University, 2024.
- 초록/해제
- 요약This thesis develops a theoretical framework for understanding the scaling properties of information processing systems in the regime of large data, large model size, and large computational resources. The goal is to develop an understanding of the impressive performance that deep neural networks have exhibited. The first part of this thesis examines models linear in their parameters but nonlinear in their inputs. This includes linear regression, kernel regression, and random feature models. Utilizing random matrix theory and free probability, I provide precise characterizations of their training dynamics, generalization capabilities, and out-of-distribution performance, alongside a detailed analysis of sources of variance. A variety of scaling laws observed in state-of-the-art large language and vision models are already present in this simple setting. The second part of this thesis focuses on representation learning. Leveraging insights from models linear in inputs but nonlinear in parameters, I present a theory of early-stage representation learning where a network with small weight initialization can learn features without altering the loss. This phenomenon, termed silent alignment, is empirically validated across various architectures and datasets. The idea of starting at small initialization leads naturally to the "maximal update parameterization", μP, that allows for feature learning at infinite width. I present empirical studies showing that practical networks can approach their theoretical infinite-width feature learning limits. Finally, I consider down-scaling the output of a neural network by a fixed constant. When this constant is small, the network behaves as a linear model in parameters; when large, it induces silent alignment. I present theoretical and empirical results of the influence of this hyperparameter on feature learning, performance, and dynamics.
- 일반주제명
- Theoretical physics
- 일반주제명
- Statistical physics
- 일반주제명
- Statistics
- 일반주제명
- Computer science
- 키워드
- Deep learning
- 기타저자
- Harvard University Physics
- 기본자료저록
- Dissertations Abstracts International. 86-05B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017164087
■00520250211152830
■006m o d
■007cr#unu||||||||
■020 ▼a9798346571490
■035 ▼a(MiAaPQ)AAI31560479
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a530.1
■1001 ▼aAtanasov, Alexander Blagoev.▼0(orcid)0000-0002-3338-0324
■24510▼aScaling and Renormalization in Statistical Learning
■260 ▼a[Sl]▼bHarvard University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a488 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-05, Section: B.
■500 ▼aAdvisor: Pehlevan, Cengiz.
■5021 ▼aThesis (Ph.D.)--Harvard University, 2024.
■520 ▼aThis thesis develops a theoretical framework for understanding the scaling properties of information processing systems in the regime of large data, large model size, and large computational resources. The goal is to develop an understanding of the impressive performance that deep neural networks have exhibited. The first part of this thesis examines models linear in their parameters but nonlinear in their inputs. This includes linear regression, kernel regression, and random feature models. Utilizing random matrix theory and free probability, I provide precise characterizations of their training dynamics, generalization capabilities, and out-of-distribution performance, alongside a detailed analysis of sources of variance. A variety of scaling laws observed in state-of-the-art large language and vision models are already present in this simple setting. The second part of this thesis focuses on representation learning. Leveraging insights from models linear in inputs but nonlinear in parameters, I present a theory of early-stage representation learning where a network with small weight initialization can learn features without altering the loss. This phenomenon, termed silent alignment, is empirically validated across various architectures and datasets. The idea of starting at small initialization leads naturally to the "maximal update parameterization", μP, that allows for feature learning at infinite width. I present empirical studies showing that practical networks can approach their theoretical infinite-width feature learning limits. Finally, I consider down-scaling the output of a neural network by a fixed constant. When this constant is small, the network behaves as a linear model in parameters; when large, it induces silent alignment. I present theoretical and empirical results of the influence of this hyperparameter on feature learning, performance, and dynamics.
■590 ▼aSchool code: 0084.
■650 4▼aTheoretical physics
■650 4▼aStatistical physics
■650 4▼aStatistics
■650 4▼aComputer science
■653 ▼aDeep learning
■653 ▼aEmpirical deep learning
■653 ▼aHigh dimensional statistics
■653 ▼aRandom matrix theory
■653 ▼aRepresentation learning
■690 ▼a0753
■690 ▼a0217
■690 ▼a0463
■690 ▼a0984
■71020▼aHarvard University▼bPhysics.
■7730 ▼tDissertations Abstracts International▼g86-05B.
■790 ▼a0084
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164087▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


