서브메뉴
검색
Representational Capabilities of Feed-Forward and Sequential Neural Architectures
Representational Capabilities of Feed-Forward and Sequential Neural Architectures
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211151155
- ISBN
- 9798382309675
- DDC
- 004
- 서명/저자
- Representational Capabilities of Feed-Forward and Sequential Neural Architectures
- 발행사항
- [Sl] : Columbia University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 414 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 85-10, Section: B.
- 주기사항
- Advisor: Hsu, Daniel J.;Servedio, Rocco A.
- 학위논문주기
- Thesis (Ph.D.)--Columbia University, 2024.
- 초록/해제
- 요약Despite the widespread empirical success of deep neural networks over the past decade, a comprehensive understanding of their mathematical properties remains elusive, which limits the abilities of practitioners to train neural networks in a principled manner. This dissertation provides a representational characterization of a variety of neural network architectures, including fully-connected feed-forward networks and sequential models like transformers. The representational capabilities of neural networks are most famously characterized by the universal approximation theorem, which states that sufficiently large neural networks can closely approximate any well-behaved target function. However, the universal approximation theorem applies exclusively to two-layer neural networks of unbounded size and fails to capture the comparative strengths and weaknesses of different architectures. The thesis addresses these limitations by quantifying the representational consequences of random features, weight regularization, and model depth on feed-forward architectures. It further investigates and contrasts the expressive powers of transformers and other sequential neural architectures. Taken together, these results apply a wide range of theoretical tools-including approximation theory, discrete dynamical systems, and communication complexity-to prove rigorous separations between different neural architectures and scaling regimes.
- 일반주제명
- Computer science
- 일반주제명
- Computer engineering
- 키워드
- Machine learning
- 키워드
- Neural networks
- 키워드
- Representation
- 키워드
- Transformers
- 기타저자
- Columbia University Computer Science
- 기본자료저록
- Dissertations Abstracts International. 85-10B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017161056
■00520250211151155
■006m o d
■007cr#unu||||||||
■020 ▼a9798382309675
■035 ▼a(MiAaPQ)AAI31236293
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aSanford, Clayton Hendrick.
■24510▼aRepresentational Capabilities of Feed-Forward and Sequential Neural Architectures
■260 ▼a[Sl]▼bColumbia University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a414 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 85-10, Section: B.
■500 ▼aAdvisor: Hsu, Daniel J.;Servedio, Rocco A.
■5021 ▼aThesis (Ph.D.)--Columbia University, 2024.
■520 ▼aDespite the widespread empirical success of deep neural networks over the past decade, a comprehensive understanding of their mathematical properties remains elusive, which limits the abilities of practitioners to train neural networks in a principled manner. This dissertation provides a representational characterization of a variety of neural network architectures, including fully-connected feed-forward networks and sequential models like transformers. The representational capabilities of neural networks are most famously characterized by the universal approximation theorem, which states that sufficiently large neural networks can closely approximate any well-behaved target function. However, the universal approximation theorem applies exclusively to two-layer neural networks of unbounded size and fails to capture the comparative strengths and weaknesses of different architectures. The thesis addresses these limitations by quantifying the representational consequences of random features, weight regularization, and model depth on feed-forward architectures. It further investigates and contrasts the expressive powers of transformers and other sequential neural architectures. Taken together, these results apply a wide range of theoretical tools-including approximation theory, discrete dynamical systems, and communication complexity-to prove rigorous separations between different neural architectures and scaling regimes.
■590 ▼aSchool code: 0054.
■650 4▼aComputer science
■650 4▼aComputer engineering
■653 ▼aMachine learning
■653 ▼aNeural networks
■653 ▼aRepresentation
■653 ▼aTheoretical computer science
■653 ▼aTransformers
■690 ▼a0984
■690 ▼a0464
■71020▼aColumbia University▼bComputer Science.
■7730 ▼tDissertations Abstracts International▼g85-10B.
■790 ▼a0054
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17161056▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


