서브메뉴
검색
Computation and Estimation for Neural Networks via Log-Concave Coupling
Computation and Estimation for Neural Networks via Log-Concave Coupling
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103031
- ISBN
- 9798286445271
- DDC
- 310
- 서명/저자
- Computation and Estimation for Neural Networks via Log-Concave Coupling
- 발행사항
- [Sl] : Yale University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 192 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-12, Section: B.
- 주기사항
- Advisor: Barron, Andrew R.
- 학위논문주기
- Thesis (Ph.D.)--Yale University, 2025.
- 초록/해제
- 요약In this work, we consider a Bayesian method to train single-hidden-layer neural networks with ℓ1 controlled weights by defining posterior distributions using different subsets of the training data, and combining posterior means to form our estimators. We consider both a joint Bayesian model for all parameters of the neural network at once, and a greedy Bayes model training the neurons one at a time based on the residuals of previous fits.The log-likelihoods of the posterior distributions we define are multimodal and nonconcave, so sampling algorithms such as Markov Chain Monte Carlo (MCMC) will not be rapidly mixing to directly sample the posteriors. Using an auxiliary random variable, we produce a mixture distribution which we call a log-concave coupling. Using a continuous uniform prior over the ℓ1 ball, the conditional distributions of this mixture are log-concave, and the mixing distribution itself is log-concave when the number of parameters in our neural network exceeds the squared number of data points. Thus the mixture distribution can be sampled efficiently to produce samples for our original target density. For a discrete uniform prior over the ℓ1 ball intersected with a grid of small spacing, we study the performance of our posterior mean estimator in an arbitrary regret sense and a statistical risk sense. Say we have a target function g, with g˜ being its projection into the closure of the convex hull of signed neurons scaled by a constant. With neuron weight vectors of dimension d and N data points, we show an estimator defined by a combination of our posterior means in the joint sampling problem has arbitrary sequence regret and statistical risk within O([(log d)/N]1/4) of the regret and risk of g˜. For the greedy construction, the additional regret and risk is an improved third root power.
- 일반주제명
- Statistics
- 일반주제명
- Mathematics
- 일반주제명
- Theoretical mathematics
- 일반주제명
- Applied mathematics
- 키워드
- Bayesian methods
- 키워드
- Machine learning
- 키워드
- Neural networks
- 기타저자
- Yale University Statistics and Data Science
- 기본자료저록
- Dissertations Abstracts International. 86-12B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017356768
■00520260202103031
■006m o d
■007cr#unu||||||||
■020 ▼a9798286445271
■035 ▼a(MiAaPQ)AAI31845940
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a310
■1001 ▼aMcDonald, Curtis James.
■24510▼aComputation and Estimation for Neural Networks via Log-Concave Coupling
■260 ▼a[Sl]▼bYale University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a192 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-12, Section: B.
■500 ▼aAdvisor: Barron, Andrew R.
■5021 ▼aThesis (Ph.D.)--Yale University, 2025.
■520 ▼aIn this work, we consider a Bayesian method to train single-hidden-layer neural networks with ℓ1 controlled weights by defining posterior distributions using different subsets of the training data, and combining posterior means to form our estimators. We consider both a joint Bayesian model for all parameters of the neural network at once, and a greedy Bayes model training the neurons one at a time based on the residuals of previous fits.The log-likelihoods of the posterior distributions we define are multimodal and nonconcave, so sampling algorithms such as Markov Chain Monte Carlo (MCMC) will not be rapidly mixing to directly sample the posteriors. Using an auxiliary random variable, we produce a mixture distribution which we call a log-concave coupling. Using a continuous uniform prior over the ℓ1 ball, the conditional distributions of this mixture are log-concave, and the mixing distribution itself is log-concave when the number of parameters in our neural network exceeds the squared number of data points. Thus the mixture distribution can be sampled efficiently to produce samples for our original target density. For a discrete uniform prior over the ℓ1 ball intersected with a grid of small spacing, we study the performance of our posterior mean estimator in an arbitrary regret sense and a statistical risk sense. Say we have a target function g, with g˜ being its projection into the closure of the convex hull of signed neurons scaled by a constant. With neuron weight vectors of dimension d and N data points, we show an estimator defined by a combination of our posterior means in the joint sampling problem has arbitrary sequence regret and statistical risk within O([(log d)/N]1/4) of the regret and risk of g˜. For the greedy construction, the additional regret and risk is an improved third root power.
■590 ▼aSchool code: 0265.
■650 4▼aStatistics
■650 4▼aMathematics
■650 4▼aTheoretical mathematics
■650 4▼aApplied mathematics
■653 ▼aBayesian methods
■653 ▼aMachine learning
■653 ▼aMarkov Chain Monte Carlo
■653 ▼aNeural networks
■653 ▼aStatistical learning theory
■690 ▼a0463
■690 ▼a0405
■690 ▼a0642
■690 ▼a0364
■71020▼aYale University▼bStatistics and Data Science.
■7730 ▼tDissertations Abstracts International▼g86-12B.
■790 ▼a0265
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17356768▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


