본문

서브메뉴

Theory of Learning in Wide Deep Neural Networks
Theory of Learning in Wide Deep Neural Networks
Theory of Learning in Wide Deep Neural Networks

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202103534
ISBN  
9798280716209
DDC  
530
저자명  
Li, Qianyi.
서명/저자  
Theory of Learning in Wide Deep Neural Networks
발행사항  
[Sl] : Harvard University, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
203 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-12, Section: B.
주기사항  
Advisor: Sompolinsky, Haim.
학위논문주기  
Thesis (Ph.D.)--Harvard University, 2025.
초록/해제  
요약In recent years, significant breakthroughs in artificial intelligence have been largely driven by advances in Deep Neural Networks (DNNs). Inspired by the layered, modular organization of the human brain, these computational models have achieved unprecedented success in fields such as image recognition, structural biology, medicine, and language. Their impressive performance is largely due to their ability to learn to extract patterns from training data and generalize these insights to previously unseen scenarios. The study of how these networks acquire knowledge - commonly known as Deep Learning Theory (DLT) - provides valuable insights into how complex networks extract meaningful features from data. Furthermore, examining the mechanisms of generalization in DNNs may yield crucial insights into how biological neural systems perform similar tasks. The remarkable generalization abilities of DNNs raise several fundamental questions: (1) How do DNNs avoid overfitting despite being heavily overparameterized? (2) What is the relation between the strength of feature learning and the DNNs' generalization capabilities? (3) DNNs commonly struggle with flexibly and continuously learning new tasks in changing environments, tasks which the human brain handles routinely and effortlessly. What factors enable or hinder DNN performance in these scenarios? To explore these questions, this dissertation introduces a framework based on a Bayesian formulation of learning, which allows us to abstract away the complexities of detailed training procedures and focus instead on the structure of the solution space. We develop a theoretical framework for wide DNNs that makes Bayesian Learning analytically tractable, and derive its main properties in learning of a single task as well as a sequence of tasks.We first analyze learning in a family of simplified, analytically tractable architectures, Deep Linear Neural Networks (DLNNs), in which each unit has a linear activation function. In the thermodynamic limit, where both the number of training examples and the width of the network become very large, yet maintain a fixed ratio, the statistics of the input-output mapping of the network, averaged throughout the solution space, can be solved exactly. Our analysis enables the evaluation of critical network properties, including generalization error, the effects of network width and depth, the size of the training set, as well as the roles of regularization and stochasticity during learning. Our theory allows for computation of both system performance and layer-wise data representations. We then heuristically extend our theory to fully connected nonlinear DNNs and validate it numerically. For a more rigorous extension to nonlinear DNNs, we propose a tractable nonlinear architecture, Globally Gated Deep Linear Networks (GGDLNs), which preserve key qualitative features of nonlinear networks while remaining analytically tractable. Compared to DLNNs, GGDLNs exhibit richer and more complex dependencies on network depth, width, and regularization. The gating operation enhances network capabilities by allowing for flexible ways to encode context. In particular, we show that GGDLNs are able to learn simultaneously multiple tasks with contradicting labels, by explicitly incorporating task-relevant information into their gating units.Finally, we extend our theoretical framework to investigate continual learning (CL) in wide DNNs, where networks sequentially learn new tasks without losing previously acquired knowledge. We first consider the single-head scenario, where a single neural network is used to perform both training and inference on all tasks. For tasks with contradicting labels which the single-head architecture struggles with, we consider a multi-head architecture with task-specific readouts. In the multi-head scenario, learning a new task involves modifying the shared hidden-layer weights while adding a new task-specific readout, leaving previous readouts untouched. This architecture can be interpreted as a gated network similar to the GGDLN, where the task identity information is incorporated into non-overlapping sets of gating units. These units then activate the corresponding output pathways for each task. Building upon the previously developed Bayesian framework, we introduce order parameters (OPs) that quantify task similarity and accurately predict the degree of forgetting and anterograde interference. Our findings emphasize that task similarity and network depth significantly impact interference in both single-head and multi-head CL setups, highlighting conditions leading to catastrophic interference and suggesting effective strategies for reducing forgetting.In summary, this dissertation presents a comprehensive theoretical analysis of learning in wide DNNs for both single and sequential tasks, demonstrating how generalization and internal representations critically depend on network architecture, hyperparameters, and task structures. These insights lay the groundwork for future exploration into generalization within more complex and practically relevant DNN architectures, and offer potential pathways toward understanding the neural mechanisms underlying representation learning and generalization in biological neural systems.
일반주제명  
Physics
일반주제명  
Neurosciences
일반주제명  
Biophysics
일반주제명  
Computational physics
키워드  
Asymptotic limits
키워드  
Continual learning
키워드  
Deep learning
키워드  
Deep Neural Networks
키워드  
Statistical mechanics
기타저자  
Harvard University Biophysics
기본자료저록  
Dissertations Abstracts International. 86-12B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017357593
■00520260202103534
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798280716209
■035    ▼a(MiAaPQ)AAI32040322
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a530
■1001  ▼aLi,  Qianyi.▼0(orcid)0000-0002-1448-4566
■24510▼aTheory  of  Learning  in  Wide  Deep  Neural  Networks
■260    ▼a[Sl]▼bHarvard  University▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a203  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-12,  Section:  B.
■500    ▼aAdvisor:  Sompolinsky,  Haim.
■5021  ▼aThesis  (Ph.D.)--Harvard  University,  2025.
■520    ▼aIn  recent  years,  significant  breakthroughs  in  artificial  intelligence  have  been  largely  driven  by  advances  in  Deep  Neural  Networks  (DNNs).  Inspired  by  the  layered,  modular  organization  of  the  human  brain,  these  computational  models  have  achieved  unprecedented  success  in  fields  such  as  image  recognition,  structural  biology,  medicine,  and  language.  Their  impressive  performance  is  largely  due  to  their  ability  to  learn  to  extract  patterns  from  training  data  and  generalize  these  insights  to  previously  unseen  scenarios.  The  study  of  how  these  networks  acquire  knowledge  -  commonly  known  as  Deep  Learning  Theory  (DLT)  -  provides  valuable  insights  into  how  complex  networks  extract  meaningful  features  from  data.  Furthermore,  examining  the  mechanisms  of  generalization  in  DNNs  may  yield  crucial  insights  into  how  biological  neural  systems  perform  similar  tasks.  The  remarkable  generalization  abilities  of  DNNs  raise  several  fundamental  questions:  (1)  How  do  DNNs  avoid  overfitting  despite  being  heavily  overparameterized?  (2)  What  is  the  relation  between  the  strength  of  feature  learning  and  the  DNNs'  generalization  capabilities?  (3)  DNNs  commonly  struggle  with  flexibly  and  continuously  learning  new  tasks  in  changing  environments,  tasks  which  the  human  brain  handles  routinely  and  effortlessly.  What  factors  enable  or  hinder  DNN  performance  in  these  scenarios?  To  explore  these  questions,  this  dissertation  introduces  a  framework  based  on  a  Bayesian  formulation  of  learning,  which  allows  us  to  abstract  away  the  complexities  of  detailed  training  procedures  and  focus  instead  on  the  structure  of  the  solution  space.  We  develop  a  theoretical  framework  for  wide  DNNs  that  makes  Bayesian  Learning  analytically  tractable,  and  derive  its  main  properties  in  learning  of  a  single  task  as  well  as  a  sequence  of  tasks.We  first  analyze  learning  in  a  family  of  simplified,  analytically  tractable  architectures,  Deep  Linear  Neural  Networks  (DLNNs),  in  which  each  unit  has  a  linear  activation  function.  In  the  thermodynamic  limit,  where  both  the  number  of  training  examples  and  the  width  of  the  network  become  very  large,  yet  maintain  a  fixed  ratio,  the  statistics  of  the  input-output  mapping  of  the  network,  averaged  throughout  the  solution  space,  can  be  solved  exactly.  Our  analysis  enables  the  evaluation  of  critical  network  properties,  including  generalization  error,  the  effects  of  network  width  and  depth,  the  size  of  the  training  set,  as  well  as  the  roles  of  regularization  and  stochasticity  during  learning.  Our  theory  allows  for  computation  of  both  system  performance  and  layer-wise  data  representations.  We  then  heuristically  extend  our  theory  to  fully  connected  nonlinear  DNNs  and  validate  it  numerically.  For  a  more  rigorous  extension  to  nonlinear  DNNs,  we  propose  a  tractable  nonlinear  architecture,  Globally  Gated  Deep  Linear  Networks  (GGDLNs),  which  preserve  key  qualitative  features  of  nonlinear  networks  while  remaining  analytically  tractable.  Compared  to  DLNNs,  GGDLNs  exhibit  richer  and  more  complex  dependencies  on  network  depth,  width,  and  regularization.  The  gating  operation  enhances  network  capabilities  by  allowing  for  flexible  ways  to  encode  context.  In  particular,  we  show  that  GGDLNs  are  able  to  learn  simultaneously  multiple  tasks  with  contradicting  labels,  by  explicitly  incorporating  task-relevant  information  into  their  gating  units.Finally,  we  extend  our  theoretical  framework  to  investigate  continual  learning  (CL)  in  wide  DNNs,  where  networks  sequentially  learn  new  tasks  without  losing  previously  acquired  knowledge.  We  first  consider  the  single-head  scenario,  where  a  single  neural  network  is  used  to  perform  both  training  and  inference  on  all  tasks.  For  tasks  with  contradicting  labels  which  the  single-head  architecture  struggles  with,  we  consider  a  multi-head  architecture  with  task-specific  readouts.  In  the  multi-head  scenario,  learning  a  new  task  involves  modifying  the  shared  hidden-layer  weights  while  adding  a  new  task-specific  readout,  leaving  previous  readouts  untouched.  This  architecture  can  be  interpreted  as  a  gated  network  similar  to  the  GGDLN,  where  the  task  identity  information  is  incorporated  into  non-overlapping  sets  of  gating  units.  These  units  then  activate  the  corresponding  output  pathways  for  each  task.  Building  upon  the  previously  developed  Bayesian  framework,  we  introduce  order  parameters  (OPs)  that  quantify  task  similarity  and  accurately  predict  the  degree  of  forgetting  and  anterograde  interference.  Our  findings  emphasize  that  task  similarity  and  network  depth  significantly  impact  interference  in  both  single-head  and  multi-head  CL  setups,  highlighting  conditions  leading  to  catastrophic  interference  and  suggesting  effective  strategies  for  reducing  forgetting.In  summary,  this  dissertation  presents  a  comprehensive  theoretical  analysis  of  learning  in  wide  DNNs  for  both  single  and  sequential  tasks,  demonstrating  how  generalization  and  internal  representations  critically  depend  on  network  architecture,  hyperparameters,  and  task  structures.  These  insights  lay  the  groundwork  for  future  exploration  into  generalization  within  more  complex  and  practically  relevant  DNN  architectures,  and  offer  potential  pathways  toward  understanding  the  neural  mechanisms  underlying  representation  learning  and  generalization  in  biological  neural  systems.
■590    ▼aSchool  code:  0084.
■650  4▼aPhysics
■650  4▼aNeurosciences
■650  4▼aBiophysics
■650  4▼aComputational  physics
■653    ▼aAsymptotic  limits
■653    ▼aContinual  learning
■653    ▼aDeep  learning
■653    ▼aDeep  Neural  Networks
■653    ▼aStatistical  mechanics
■690    ▼a0605
■690    ▼a0317
■690    ▼a0786
■690    ▼a0216
■71020▼aHarvard  University▼bBiophysics.
■7730  ▼tDissertations  Abstracts  International▼g86-12B.
■790    ▼a0084
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357593▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF17961 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.