서브메뉴
검색
Convex Optimization Formulation of Neural Networks: Theories, Applications and Beyond
Convex Optimization Formulation of Neural Networks: Theories, Applications and Beyond
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103136
- ISBN
- 9798311950862
- DDC
- 530
- 저자명
- Wang, Yifei.
- 서명/저자
- Convex Optimization Formulation of Neural Networks: Theories, Applications and Beyond
- 발행사항
- [Sl] : Stanford University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 372 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-12, Section: B.
- 주기사항
- Advisor: Pilanci, Mert.
- 학위논문주기
- Thesis (Ph.D.)--Stanford University, 2025.
- 초록/해제
- 요약Deep neural networks (DNNs) have revolutionized numerous fields, including computer vision, natural language processing (NLP), and recommendation systems, demonstrating remarkable capabilities in representation learning and generalization. Their empirical success is driven by highly expressive architectures, large-scale datasets, and sophisticated training strategies. However, despite these advancements, a complete theoretical understanding of their optimization and generalization properties remains an open challenge. The intrinsic nonlinearity of neural networks, coupled with over-parameterization and the highly nonconvex nature of their training landscapes, poses significant difficulties in theoretical analysis.Large language models (LLMs), a specialized class of deep networks, have further pushed the boundaries of artificial intelligence by achieving unprecedented performance across a wide range of tasks. However, training such models is computationally intensive, requiring massive datasets and substantial computational resources. Even fine-tuning pretrained LLMs for specific tasks remains a costly and resource-demanding process, limiting their accessibility and scalability. These challenges underscore the need for alternative optimization frameworks that can enhance training efficiency without compromising performance.One promising avenue in this regard is convex optimization, which offers a principled approach to designing efficient training algorithms for neural networks. By leveraging convex formulations, it may be possible to mitigate the computational challenges associated with deep learning while preserving the expressiveness and generalization capabilities of neural networks. This thesis explores how convex optimization techniques can provide new theoretical insights and practical improvements in neural network training, paving the way for more efficient and scalable learning paradigms.Let X ∈ R nxdand y ∈ R nbe the data matrix and the label vector. Given a number of neurons m⩾ 1 and a regularization parameter β 0, we consider the regularized optimization problem of the simplist neural network, a two-layer ReLU neural network:where Θm= R dxmx R m, θ = (W1, W2), w1,iis the i-th column of W1∈ R dxmand w2,iis the i-th coefficient of w2∈ R m. Here we focus on the ReLU activation, i.e., σ(z) = max{z, 0} and absorb the label y ∈ R nin the loss function ℓ : R n→ R, which is assumed to be convex (e.g., logistic, hinge, squared loss). The model Σmi=1σ(Xw1,i)w2,iin (1.1) can be easily extended to the one with bias term by adding a column of 1's into the data X. We refer to an element θ ∈ Θmas a neural network and to each pair (w1,i, w2,i) as a neuron. We denote the set of optimal neural network as We denote the best training loss achievable by a neural as P* = infm⩾1 P ∗ m.We now introduce an important concept from combinatorial geometry called hyperplane arrangement patterns, which plays an important role in convex optimization formulations of ReLU network training problems.
- 일반주제명
- Phase transitions
- 일반주제명
- Convex analysis
- 일반주제명
- Neurons
- 일반주제명
- Algebra
- 일반주제명
- Neural networks
- 일반주제명
- Computer science
- 기타저자
- Stanford University.
- 기본자료저록
- Dissertations Abstracts International. 86-12B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017357129
■00520260202103136
■006m o d
■007cr#unu||||||||
■020 ▼a9798311950862
■035 ▼a(MiAaPQ)AAI31974609
■035 ▼a(MiAaPQ)Stanfordhh655zv7345
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a530
■1001 ▼aWang, Yifei.
■24510▼aConvex Optimization Formulation of Neural Networks: Theories, Applications and Beyond
■260 ▼a[Sl]▼bStanford University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a372 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-12, Section: B.
■500 ▼aAdvisor: Pilanci, Mert.
■5021 ▼aThesis (Ph.D.)--Stanford University, 2025.
■520 ▼aDeep neural networks (DNNs) have revolutionized numerous fields, including computer vision, natural language processing (NLP), and recommendation systems, demonstrating remarkable capabilities in representation learning and generalization. Their empirical success is driven by highly expressive architectures, large-scale datasets, and sophisticated training strategies. However, despite these advancements, a complete theoretical understanding of their optimization and generalization properties remains an open challenge. The intrinsic nonlinearity of neural networks, coupled with over-parameterization and the highly nonconvex nature of their training landscapes, poses significant difficulties in theoretical analysis.Large language models (LLMs), a specialized class of deep networks, have further pushed the boundaries of artificial intelligence by achieving unprecedented performance across a wide range of tasks. However, training such models is computationally intensive, requiring massive datasets and substantial computational resources. Even fine-tuning pretrained LLMs for specific tasks remains a costly and resource-demanding process, limiting their accessibility and scalability. These challenges underscore the need for alternative optimization frameworks that can enhance training efficiency without compromising performance.One promising avenue in this regard is convex optimization, which offers a principled approach to designing efficient training algorithms for neural networks. By leveraging convex formulations, it may be possible to mitigate the computational challenges associated with deep learning while preserving the expressiveness and generalization capabilities of neural networks. This thesis explores how convex optimization techniques can provide new theoretical insights and practical improvements in neural network training, paving the way for more efficient and scalable learning paradigms.Let X ∈ R nxdand y ∈ R nbe the data matrix and the label vector. Given a number of neurons m⩾ 1 and a regularization parameter β 0, we consider the regularized optimization problem of the simplist neural network, a two-layer ReLU neural network:where Θm= R dxmx R m, θ = (W1, W2), w1,iis the i-th column of W1∈ R dxmand w2,iis the i-th coefficient of w2∈ R m. Here we focus on the ReLU activation, i.e., σ(z) = max{z, 0} and absorb the label y ∈ R nin the loss function ℓ : R n→ R, which is assumed to be convex (e.g., logistic, hinge, squared loss). The model Σmi=1σ(Xw1,i)w2,iin (1.1) can be easily extended to the one with bias term by adding a column of 1's into the data X. We refer to an element θ ∈ Θmas a neural network and to each pair (w1,i, w2,i) as a neuron. We denote the set of optimal neural network as We denote the best training loss achievable by a neural as P* = infm⩾1 P ∗ m.We now introduce an important concept from combinatorial geometry called hyperplane arrangement patterns, which plays an important role in convex optimization formulations of ReLU network training problems.
■590 ▼aSchool code: 0212.
■650 4▼aPhase transitions
■650 4▼aConvex analysis
■650 4▼aNeurons
■650 4▼aAlgebra
■650 4▼aNeural networks
■650 4▼aComputer science
■690 ▼a0984
■690 ▼a0800
■71020▼aStanford University.
■7730 ▼tDissertations Abstracts International▼g86-12B.
■790 ▼a0212
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357129▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


