본문

서브메뉴

Science of Deep Learning: From Initialization to Emergent Structures
Science of Deep Learning: From Initialization to Emergent Structures
Science of Deep Learning: From Initialization to Emergent Structures

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202103124
ISBN  
9798286436750
DDC  
530
저자명  
Doshi, Darshil.
서명/저자  
Science of Deep Learning: From Initialization to Emergent Structures
발행사항  
[Sl] : University of Maryland, College Park, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
262 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-12, Section: A.
주기사항  
Advisor: Barkeshli, Maissam;Gromov, Andrey.
학위논문주기  
Thesis (Ph.D.)--University of Maryland, College Park, 2025.
초록/해제  
요약As artificial intelligence (AI) systems grow increasingly powerful and permeate every aspect of our lives, their impact on both individuals and society is an urgent concern. Questions of safety and robustness in AI stem largely from our limited understanding of deep learning. Research in this domain has traditionally followed two parallel paths: an empirical approach that prioritizes practical advancements and a theoretical approach that seeks a mathematical understanding from first principles. Despite notable progress, a significant gap remains between deep learning practice and its theoretical underpinnings. This dissertation advocates for a phenomenological approach to understanding AI systems -- one that integrates empirical observations with theoretical model-building. This methodology has been instrumental in the physical sciences, and it holds similar promise for advancing the science of deep learning. Over two broad parts, this work demonstrates the effectiveness of this approach in characterizing model architectures and their emergent capabilities.In the first part, we explore how signal propagation analysis in large-N limits can inform the design and initialization of model architectures. We develop a diagnostic observable that distinguishes between ordered and chaotic behaviors in neural networks, guiding optimal parameter initialization for training. Our analysis establishes the theoretical soundness of this observable in simple networks and confirms its empirical utility in state-of-the-art architectures. The findings reveal an architecture design paradigm that eliminates the need for careful initialization, shedding light on widely used heuristic practices. Additionally, we introduce an algorithm that automates initialization across diverse model architectures, enhancing their trainability.In the second part, we highlight the importance of the systems identification approach for characterizing AI systems. We explore several stylized setups where model capabilities emerge as a function of compute, data quantity, and data diversity. Using arithmetic and cryptographic tasks as examples, we demonstrate that emergent abilities such as grokking and in-context learning arise alongside the formation of interpretable structures within the model's parameters, hidden representations, and outputs. Through targeted experiments, we identify these structures using (i) black-box probing, which examines model responses to characteristic inputs, and (ii) open-box analysis, which leverages curated task-specific observables and metrics to study internal model states.This dissertation promotes a paradigm for understanding deep learning that complements both heuristic-driven and hypothesis-driven approaches. By integrating experimental methodologies and analytical tools from established scientific disciplines, this framework has the potential to steer the field toward safer, more robust, and more efficient AI systems.
일반주제명  
Physics
일반주제명  
Information science
키워드  
AI interpretability
키워드  
Critical initialization
키워드  
Deep learning
키워드  
Emergence
키워드  
Grokking
키워드  
In-context learning
기타저자  
University of Maryland, College Park Physics
기본자료저록  
Dissertations Abstracts International. 86-12A.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017357056
■00520260202103124
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798286436750
■035    ▼a(MiAaPQ)AAI31938513
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a530
■1001  ▼aDoshi,  Darshil.▼0(orcid)0000-0003-3578-9016
■24510▼aScience  of  Deep  Learning:  From  Initialization  to  Emergent  Structures
■260    ▼a[Sl]▼bUniversity  of  Maryland,  College  Park▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a262  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-12,  Section:  A.
■500    ▼aAdvisor:  Barkeshli,  Maissam;Gromov,  Andrey.
■5021  ▼aThesis  (Ph.D.)--University  of  Maryland,  College  Park,  2025.
■520    ▼aAs  artificial  intelligence  (AI)  systems  grow  increasingly  powerful  and  permeate  every  aspect  of  our  lives,  their  impact  on  both  individuals  and  society  is  an  urgent  concern.  Questions  of  safety  and  robustness  in  AI  stem  largely  from  our  limited  understanding  of  deep  learning.  Research  in  this  domain  has  traditionally  followed  two  parallel  paths:  an  empirical  approach  that  prioritizes  practical  advancements  and  a  theoretical  approach  that  seeks  a  mathematical  understanding  from  first  principles.  Despite  notable  progress,  a  significant  gap  remains  between  deep  learning  practice  and  its  theoretical  underpinnings.  This  dissertation  advocates  for  a  phenomenological  approach  to  understanding  AI  systems  --  one  that  integrates  empirical  observations  with  theoretical  model-building.  This  methodology  has  been  instrumental  in  the  physical  sciences,  and  it  holds  similar  promise  for  advancing  the  science  of  deep  learning.  Over  two  broad  parts,  this  work  demonstrates  the  effectiveness  of  this  approach  in  characterizing  model  architectures  and  their  emergent  capabilities.In  the  first  part,  we  explore  how  signal  propagation  analysis  in  large-N  limits  can  inform  the  design  and  initialization  of  model  architectures.  We  develop  a  diagnostic  observable  that  distinguishes  between  ordered  and  chaotic  behaviors  in  neural  networks,  guiding  optimal  parameter  initialization  for  training.  Our  analysis  establishes  the  theoretical  soundness  of  this  observable  in  simple  networks  and  confirms  its  empirical  utility  in  state-of-the-art  architectures.  The  findings  reveal  an  architecture  design  paradigm  that  eliminates  the  need  for  careful  initialization,  shedding  light  on  widely  used  heuristic  practices.  Additionally,  we  introduce  an  algorithm  that  automates  initialization  across  diverse  model  architectures,  enhancing  their  trainability.In  the  second  part,  we  highlight  the  importance  of  the  systems  identification  approach  for  characterizing  AI  systems.  We  explore  several  stylized  setups  where  model  capabilities  emerge  as  a  function  of  compute,  data  quantity,  and  data  diversity.  Using  arithmetic  and  cryptographic  tasks  as  examples,  we  demonstrate  that  emergent  abilities  such  as  grokking  and  in-context  learning  arise  alongside  the  formation  of  interpretable  structures  within  the  model's  parameters,  hidden  representations,  and  outputs.  Through  targeted  experiments,  we  identify  these  structures  using  (i)  black-box  probing,  which  examines  model  responses  to  characteristic  inputs,  and  (ii)  open-box  analysis,  which  leverages  curated  task-specific  observables  and  metrics  to  study  internal  model  states.This  dissertation  promotes  a  paradigm  for  understanding  deep  learning  that  complements  both  heuristic-driven  and  hypothesis-driven  approaches.  By  integrating  experimental  methodologies  and  analytical  tools  from  established  scientific  disciplines,  this  framework  has  the  potential  to  steer  the  field  toward  safer,  more  robust,  and  more  efficient  AI  systems.
■590    ▼aSchool  code:  0117.
■650  4▼aPhysics
■650  4▼aInformation  science
■653    ▼aAI  interpretability
■653    ▼aCritical  initialization
■653    ▼aDeep  learning
■653    ▼aEmergence
■653    ▼aGrokking
■653    ▼aIn-context  learning
■690    ▼a0605
■690    ▼a0800
■690    ▼a0723
■71020▼aUniversity  of  Maryland,  College  Park▼bPhysics.
■7730  ▼tDissertations  Abstracts  International▼g86-12A.
■790    ▼a0117
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357056▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF15382 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.