본문

서브메뉴

Convex Optimization Formulation of Neural Networks: Theories, Applications and Beyond
Convex Optimization Formulation of Neural Networks: Theories, Applications and Beyond
Convex Optimization Formulation of Neural Networks: Theories, Applications and Beyond

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202103136
ISBN  
9798311950862
DDC  
530
저자명  
Wang, Yifei.
서명/저자  
Convex Optimization Formulation of Neural Networks: Theories, Applications and Beyond
발행사항  
[Sl] : Stanford University, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
372 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-12, Section: B.
주기사항  
Advisor: Pilanci, Mert.
학위논문주기  
Thesis (Ph.D.)--Stanford University, 2025.
초록/해제  
요약Deep neural networks (DNNs) have revolutionized numerous fields, including computer vision, natural language processing (NLP), and recommendation systems, demonstrating remarkable capabilities in representation learning and generalization. Their empirical success is driven by highly expressive architectures, large-scale datasets, and sophisticated training strategies. However, despite these advancements, a complete theoretical understanding of their optimization and generalization properties remains an open challenge. The intrinsic nonlinearity of neural networks, coupled with over-parameterization and the highly nonconvex nature of their training landscapes, poses significant difficulties in theoretical analysis.Large language models (LLMs), a specialized class of deep networks, have further pushed the boundaries of artificial intelligence by achieving unprecedented performance across a wide range of tasks. However, training such models is computationally intensive, requiring massive datasets and substantial computational resources. Even fine-tuning pretrained LLMs for specific tasks remains a costly and resource-demanding process, limiting their accessibility and scalability. These challenges underscore the need for alternative optimization frameworks that can enhance training efficiency without compromising performance.One promising avenue in this regard is convex optimization, which offers a principled approach to designing efficient training algorithms for neural networks. By leveraging convex formulations, it may be possible to mitigate the computational challenges associated with deep learning while preserving the expressiveness and generalization capabilities of neural networks. This thesis explores how convex optimization techniques can provide new theoretical insights and practical improvements in neural network training, paving the way for more efficient and scalable learning paradigms.Let X ∈ R nxdand y ∈ R nbe the data matrix and the label vector. Given a number of neurons m⩾ 1 and a regularization parameter β 0, we consider the regularized optimization problem of the simplist neural network, a two-layer ReLU neural network:where Θm= R dxmx R m, θ = (W1, W2), w1,iis the i-th column of W1∈ R dxmand w2,iis the i-th coefficient of w2∈ R m. Here we focus on the ReLU activation, i.e., σ(z) = max{z, 0} and absorb the label y ∈ R nin the loss function ℓ : R n→ R, which is assumed to be convex (e.g., logistic, hinge, squared loss). The model Σmi=1σ(Xw1,i)w2,iin (1.1) can be easily extended to the one with bias term by adding a column of 1's into the data X. We refer to an element θ ∈ Θmas a neural network and to each pair (w1,i, w2,i) as a neuron. We denote the set of optimal neural network as We denote the best training loss achievable by a neural as P* = infm⩾1 P ∗ m.We now introduce an important concept from combinatorial geometry called hyperplane arrangement patterns, which plays an important role in convex optimization formulations of ReLU network training problems.
일반주제명  
Phase transitions
일반주제명  
Convex analysis
일반주제명  
Neurons
일반주제명  
Algebra
일반주제명  
Neural networks
일반주제명  
Computer science
기타저자  
Stanford University.
기본자료저록  
Dissertations Abstracts International. 86-12B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017357129
■00520260202103136
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798311950862
■035    ▼a(MiAaPQ)AAI31974609
■035    ▼a(MiAaPQ)Stanfordhh655zv7345
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a530
■1001  ▼aWang,  Yifei.
■24510▼aConvex  Optimization  Formulation  of  Neural  Networks:  Theories,  Applications  and  Beyond
■260    ▼a[Sl]▼bStanford  University▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a372  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-12,  Section:  B.
■500    ▼aAdvisor:  Pilanci,  Mert.
■5021  ▼aThesis  (Ph.D.)--Stanford  University,  2025.
■520    ▼aDeep  neural  networks  (DNNs)  have  revolutionized  numerous  fields,  including  computer  vision,  natural  language  processing  (NLP),  and  recommendation  systems,  demonstrating  remarkable  capabilities  in  representation  learning  and  generalization.  Their  empirical  success  is  driven  by  highly  expressive  architectures,  large-scale  datasets,  and  sophisticated  training  strategies.  However,  despite  these  advancements,  a  complete  theoretical  understanding  of  their  optimization  and  generalization  properties  remains  an  open  challenge.  The  intrinsic  nonlinearity  of  neural  networks,  coupled  with  over-parameterization  and  the  highly  nonconvex  nature  of  their  training  landscapes,  poses  significant  difficulties  in  theoretical  analysis.Large  language  models  (LLMs),  a  specialized  class  of  deep  networks,  have  further  pushed  the  boundaries  of  artificial  intelligence  by  achieving  unprecedented  performance  across  a  wide  range  of  tasks.  However,  training  such  models  is  computationally  intensive,  requiring  massive  datasets  and  substantial  computational  resources.  Even  fine-tuning  pretrained  LLMs  for  specific  tasks  remains  a  costly  and  resource-demanding  process,  limiting  their  accessibility  and  scalability.  These  challenges  underscore  the  need  for  alternative  optimization  frameworks  that  can  enhance  training  efficiency  without  compromising  performance.One  promising  avenue  in  this  regard  is  convex  optimization,  which  offers  a  principled  approach  to  designing  efficient  training  algorithms  for  neural  networks.  By  leveraging  convex  formulations,  it  may  be  possible  to  mitigate  the  computational  challenges  associated  with  deep  learning  while  preserving  the  expressiveness  and  generalization  capabilities  of  neural  networks.  This  thesis  explores  how  convex  optimization  techniques  can  provide  new  theoretical  insights  and  practical  improvements  in  neural  network  training,  paving  the  way  for  more  efficient  and  scalable  learning  paradigms.Let  X  ∈  R  nxdand  y  ∈  R  nbe  the  data  matrix  and  the  label  vector.  Given  a  number  of  neurons  m⩾  1  and  a  regularization  parameter  β  0,  we  consider  the  regularized  optimization  problem  of  the  simplist  neural  network,  a  two-layer  ReLU  neural  network:where  Θm=  R  dxmx  R  m,  θ  =  (W1,  W2),  w1,iis  the  i-th  column  of  W1∈  R  dxmand  w2,iis  the  i-th  coefficient  of  w2∈  R  m.  Here  we  focus  on  the  ReLU  activation,  i.e.,  σ(z)  =  max{z,  0}  and  absorb  the  label  y  ∈  R  nin  the  loss  function  ℓ  :  R  n→  R,  which  is  assumed  to  be  convex  (e.g.,  logistic,  hinge,  squared  loss).  The  model  Σmi=1σ(Xw1,i)w2,iin  (1.1)  can  be  easily  extended  to  the  one  with  bias  term  by  adding  a  column  of  1's  into  the  data  X.  We  refer  to  an  element  θ  ∈  Θmas  a  neural  network  and  to  each  pair  (w1,i,  w2,i)  as  a  neuron.  We  denote  the  set  of  optimal  neural  network  as  We  denote  the  best  training  loss  achievable  by  a  neural  as  P*  =  infm⩾1  P  ∗  m.We  now  introduce  an  important  concept  from  combinatorial  geometry  called  hyperplane  arrangement  patterns,  which  plays  an  important  role  in  convex  optimization  formulations  of  ReLU  network  training  problems.
■590    ▼aSchool  code:  0212.
■650  4▼aPhase  transitions
■650  4▼aConvex  analysis
■650  4▼aNeurons
■650  4▼aAlgebra
■650  4▼aNeural  networks
■650  4▼aComputer  science
■690    ▼a0984
■690    ▼a0800
■71020▼aStanford  University.
■7730  ▼tDissertations  Abstracts  International▼g86-12B.
■790    ▼a0212
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357129▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF16806 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.