서브메뉴
검색
Science of Deep Learning: From Initialization to Emergent Structures
Science of Deep Learning: From Initialization to Emergent Structures
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103124
- ISBN
- 9798286436750
- DDC
- 530
- 저자명
- Doshi, Darshil.
- 서명/저자
- Science of Deep Learning: From Initialization to Emergent Structures
- 발행사항
- [Sl] : University of Maryland, College Park, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 262 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-12, Section: A.
- 주기사항
- Advisor: Barkeshli, Maissam;Gromov, Andrey.
- 학위논문주기
- Thesis (Ph.D.)--University of Maryland, College Park, 2025.
- 초록/해제
- 요약As artificial intelligence (AI) systems grow increasingly powerful and permeate every aspect of our lives, their impact on both individuals and society is an urgent concern. Questions of safety and robustness in AI stem largely from our limited understanding of deep learning. Research in this domain has traditionally followed two parallel paths: an empirical approach that prioritizes practical advancements and a theoretical approach that seeks a mathematical understanding from first principles. Despite notable progress, a significant gap remains between deep learning practice and its theoretical underpinnings. This dissertation advocates for a phenomenological approach to understanding AI systems -- one that integrates empirical observations with theoretical model-building. This methodology has been instrumental in the physical sciences, and it holds similar promise for advancing the science of deep learning. Over two broad parts, this work demonstrates the effectiveness of this approach in characterizing model architectures and their emergent capabilities.In the first part, we explore how signal propagation analysis in large-N limits can inform the design and initialization of model architectures. We develop a diagnostic observable that distinguishes between ordered and chaotic behaviors in neural networks, guiding optimal parameter initialization for training. Our analysis establishes the theoretical soundness of this observable in simple networks and confirms its empirical utility in state-of-the-art architectures. The findings reveal an architecture design paradigm that eliminates the need for careful initialization, shedding light on widely used heuristic practices. Additionally, we introduce an algorithm that automates initialization across diverse model architectures, enhancing their trainability.In the second part, we highlight the importance of the systems identification approach for characterizing AI systems. We explore several stylized setups where model capabilities emerge as a function of compute, data quantity, and data diversity. Using arithmetic and cryptographic tasks as examples, we demonstrate that emergent abilities such as grokking and in-context learning arise alongside the formation of interpretable structures within the model's parameters, hidden representations, and outputs. Through targeted experiments, we identify these structures using (i) black-box probing, which examines model responses to characteristic inputs, and (ii) open-box analysis, which leverages curated task-specific observables and metrics to study internal model states.This dissertation promotes a paradigm for understanding deep learning that complements both heuristic-driven and hypothesis-driven approaches. By integrating experimental methodologies and analytical tools from established scientific disciplines, this framework has the potential to steer the field toward safer, more robust, and more efficient AI systems.
- 일반주제명
- Physics
- 일반주제명
- Information science
- 키워드
- Deep learning
- 키워드
- Emergence
- 키워드
- Grokking
- 기타저자
- University of Maryland, College Park Physics
- 기본자료저록
- Dissertations Abstracts International. 86-12A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017357056
■00520260202103124
■006m o d
■007cr#unu||||||||
■020 ▼a9798286436750
■035 ▼a(MiAaPQ)AAI31938513
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a530
■1001 ▼aDoshi, Darshil.▼0(orcid)0000-0003-3578-9016
■24510▼aScience of Deep Learning: From Initialization to Emergent Structures
■260 ▼a[Sl]▼bUniversity of Maryland, College Park▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a262 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-12, Section: A.
■500 ▼aAdvisor: Barkeshli, Maissam;Gromov, Andrey.
■5021 ▼aThesis (Ph.D.)--University of Maryland, College Park, 2025.
■520 ▼aAs artificial intelligence (AI) systems grow increasingly powerful and permeate every aspect of our lives, their impact on both individuals and society is an urgent concern. Questions of safety and robustness in AI stem largely from our limited understanding of deep learning. Research in this domain has traditionally followed two parallel paths: an empirical approach that prioritizes practical advancements and a theoretical approach that seeks a mathematical understanding from first principles. Despite notable progress, a significant gap remains between deep learning practice and its theoretical underpinnings. This dissertation advocates for a phenomenological approach to understanding AI systems -- one that integrates empirical observations with theoretical model-building. This methodology has been instrumental in the physical sciences, and it holds similar promise for advancing the science of deep learning. Over two broad parts, this work demonstrates the effectiveness of this approach in characterizing model architectures and their emergent capabilities.In the first part, we explore how signal propagation analysis in large-N limits can inform the design and initialization of model architectures. We develop a diagnostic observable that distinguishes between ordered and chaotic behaviors in neural networks, guiding optimal parameter initialization for training. Our analysis establishes the theoretical soundness of this observable in simple networks and confirms its empirical utility in state-of-the-art architectures. The findings reveal an architecture design paradigm that eliminates the need for careful initialization, shedding light on widely used heuristic practices. Additionally, we introduce an algorithm that automates initialization across diverse model architectures, enhancing their trainability.In the second part, we highlight the importance of the systems identification approach for characterizing AI systems. We explore several stylized setups where model capabilities emerge as a function of compute, data quantity, and data diversity. Using arithmetic and cryptographic tasks as examples, we demonstrate that emergent abilities such as grokking and in-context learning arise alongside the formation of interpretable structures within the model's parameters, hidden representations, and outputs. Through targeted experiments, we identify these structures using (i) black-box probing, which examines model responses to characteristic inputs, and (ii) open-box analysis, which leverages curated task-specific observables and metrics to study internal model states.This dissertation promotes a paradigm for understanding deep learning that complements both heuristic-driven and hypothesis-driven approaches. By integrating experimental methodologies and analytical tools from established scientific disciplines, this framework has the potential to steer the field toward safer, more robust, and more efficient AI systems.
■590 ▼aSchool code: 0117.
■650 4▼aPhysics
■650 4▼aInformation science
■653 ▼aAI interpretability
■653 ▼aCritical initialization
■653 ▼aDeep learning
■653 ▼aEmergence
■653 ▼aGrokking
■653 ▼aIn-context learning
■690 ▼a0605
■690 ▼a0800
■690 ▼a0723
■71020▼aUniversity of Maryland, College Park▼bPhysics.
■7730 ▼tDissertations Abstracts International▼g86-12A.
■790 ▼a0117
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357056▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


