본문

서브메뉴

Scaling and Renormalization in Statistical Learning
Scaling and Renormalization in Statistical Learning
Scaling and Renormalization in Statistical Learning

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211152830
ISBN  
9798346571490
DDC  
530.1
저자명  
Atanasov, Alexander Blagoev.
서명/저자  
Scaling and Renormalization in Statistical Learning
발행사항  
[Sl] : Harvard University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
488 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-05, Section: B.
주기사항  
Advisor: Pehlevan, Cengiz.
학위논문주기  
Thesis (Ph.D.)--Harvard University, 2024.
초록/해제  
요약This thesis develops a theoretical framework for understanding the scaling properties of information processing systems in the regime of large data, large model size, and large computational resources. The goal is to develop an understanding of the impressive performance that deep neural networks have exhibited. The first part of this thesis examines models linear in their parameters but nonlinear in their inputs. This includes linear regression, kernel regression, and random feature models. Utilizing random matrix theory and free probability, I provide precise characterizations of their training dynamics, generalization capabilities, and out-of-distribution performance, alongside a detailed analysis of sources of variance. A variety of scaling laws observed in state-of-the-art large language and vision models are already present in this simple setting. The second part of this thesis focuses on representation learning. Leveraging insights from models linear in inputs but nonlinear in parameters, I present a theory of early-stage representation learning where a network with small weight initialization can learn features without altering the loss. This phenomenon, termed silent alignment, is empirically validated across various architectures and datasets. The idea of starting at small initialization leads naturally to the "maximal update parameterization", μP, that allows for feature learning at infinite width. I present empirical studies showing that practical networks can approach their theoretical infinite-width feature learning limits. Finally, I consider down-scaling the output of a neural network by a fixed constant. When this constant is small, the network behaves as a linear model in parameters; when large, it induces silent alignment. I present theoretical and empirical results of the influence of this hyperparameter on feature learning, performance, and dynamics.
일반주제명  
Theoretical physics
일반주제명  
Statistical physics
일반주제명  
Statistics
일반주제명  
Computer science
키워드  
Deep learning
키워드  
Empirical deep learning
키워드  
High dimensional statistics
키워드  
Random matrix theory
키워드  
Representation learning
기타저자  
Harvard University Physics
기본자료저록  
Dissertations Abstracts International. 86-05B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017164087
■00520250211152830
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798346571490
■035    ▼a(MiAaPQ)AAI31560479
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a530.1
■1001  ▼aAtanasov,  Alexander  Blagoev.▼0(orcid)0000-0002-3338-0324
■24510▼aScaling  and  Renormalization  in  Statistical  Learning
■260    ▼a[Sl]▼bHarvard  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a488  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-05,  Section:  B.
■500    ▼aAdvisor:  Pehlevan,  Cengiz.
■5021  ▼aThesis  (Ph.D.)--Harvard  University,  2024.
■520    ▼aThis  thesis  develops  a  theoretical  framework  for  understanding  the  scaling  properties  of  information  processing  systems  in  the  regime  of  large  data,  large  model  size,  and  large  computational  resources.  The  goal  is  to  develop  an  understanding  of  the  impressive  performance  that  deep  neural  networks  have  exhibited.  The  first  part  of  this  thesis  examines  models  linear  in  their  parameters  but  nonlinear  in  their  inputs.  This  includes  linear  regression,  kernel  regression,  and  random  feature  models.  Utilizing  random  matrix  theory  and  free  probability,  I  provide  precise  characterizations  of  their  training  dynamics,  generalization  capabilities,  and  out-of-distribution  performance,  alongside  a  detailed  analysis  of  sources  of  variance.  A  variety  of  scaling  laws  observed  in  state-of-the-art  large  language  and  vision  models  are  already  present  in  this  simple  setting.  The  second  part  of  this  thesis  focuses  on  representation  learning.  Leveraging  insights  from  models  linear  in  inputs  but  nonlinear  in  parameters,  I  present  a  theory  of  early-stage  representation  learning  where  a  network  with  small  weight  initialization  can  learn  features  without  altering  the  loss.  This  phenomenon,  termed  silent  alignment,  is  empirically  validated  across  various  architectures  and  datasets.  The  idea  of  starting  at  small  initialization  leads  naturally  to  the  "maximal  update  parameterization",  μP,  that  allows  for  feature  learning  at  infinite  width.  I  present  empirical  studies  showing  that  practical  networks  can  approach  their  theoretical  infinite-width  feature  learning  limits.  Finally,  I  consider  down-scaling  the  output  of  a  neural  network  by  a  fixed  constant.  When  this  constant  is  small,  the  network  behaves  as  a  linear  model  in  parameters;  when  large,  it  induces  silent  alignment.  I  present  theoretical  and  empirical  results  of  the  influence  of  this  hyperparameter  on  feature  learning,  performance,  and  dynamics. 
■590    ▼aSchool  code:  0084.
■650  4▼aTheoretical  physics
■650  4▼aStatistical  physics
■650  4▼aStatistics
■650  4▼aComputer  science
■653    ▼aDeep  learning
■653    ▼aEmpirical  deep  learning
■653    ▼aHigh  dimensional  statistics
■653    ▼aRandom  matrix  theory
■653    ▼aRepresentation  learning
■690    ▼a0753
■690    ▼a0217
■690    ▼a0463
■690    ▼a0984
■71020▼aHarvard  University▼bPhysics.
■7730  ▼tDissertations  Abstracts  International▼g86-05B.
■790    ▼a0084
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17164087▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF14335 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.