서브메뉴
검색
Trade-Offs and Opportunities in High-Dimensional Bayesian Modeling
Trade-Offs and Opportunities in High-Dimensional Bayesian Modeling
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211152029
- ISBN
- 9798383705735
- DDC
- 310
- 서명/저자
- Trade-Offs and Opportunities in High-Dimensional Bayesian Modeling
- 발행사항
- [Sl] : Columbia University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 259 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-02, Section: A.
- 주기사항
- Advisor: Gelman, Andrew;Rush, Cynthia.
- 학위논문주기
- Thesis (Ph.D.)--Columbia University, 2024.
- 초록/해제
- 요약With the increasing availability of large multivariate datasets, modern parametric statistical models makes increasing use of high-dimensional parameter spaces to flexibly represent complex data generating mechanisms. Yet, ceteris paribus, increases in dimensionality often carry drawbacks across the various sub-problems of data analysis, posing challenges for the data analyst who must balance model plausibility against the practical considerations of implementation. We focus here on challenges to three components of data analysis: computation, inference, and model checking. In the computational domain, we are concerned with achieving reasonable scaling of the computational complexity with the parameter dimension without sacrificing the trustworthiness of our computation. Here, we study a particular class of algorithms - the vectorized approximate message passing (VAMP) iterations - which offer the possibility of linear per-iteration scaling with dimension. These iterations perform approximate inference for a class of Bayesian generalized linear regression models, and we demonstrate that under flexible distributional conditions, the estimation performance of these VAMP iterations can be predicted to high accuracy with probability decaying exponentially fast in the size of the regression problem. In the realm of statistical inference, we investigate the relationship between parameter dimension and identification. We develop formal notions of weak identification and model expansion in the Bayesian setting and use this to argue for a very general tendency for dimensionality-increasing model expansion to weaken the identification of model parameters. We draw two substantive conclusions from this formalism. First, the negative association between dimensionality and identification can be weakened or reversed when we construct prior distributions that encode sufficiently strong dependence between parameters. Absent such prior information, we derive bounds which indicate that decreasing identification is usually unavoidable with sufficient inflation of the dimension without increasing the severity of the third challenge we consider: that of dimensionality to model checking.We divide the topic of model checking into two sub-problems: fitness testing and correctness testing. Using our model expansion formalism, we show again that both of these problems tend to become more difficult as the model dimension grows. We propose two extensions of the posterior predictive \uD835\uDC5D-value - certain conditional and joint \uD835\uDC5D-values, which are designed to address these challenges for fitness and correctness testing respectively. We demonstrate the potential of these \uD835\uDC5D-values to allow successful model checking that scales with dimensionality theoretically and with examples.
- 일반주제명
- Statistics
- 일반주제명
- Statistical physics
- 일반주제명
- Information science
- 키워드
- Model checking
- 기타저자
- Columbia University Statistics
- 기본자료저록
- Dissertations Abstracts International. 86-02A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017162579
■00520250211152029
■006m o d
■007cr#unu||||||||
■020 ▼a9798383705735
■035 ▼a(MiAaPQ)AAI31333836
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a310
■1001 ▼aCademartori, Collin Andrew.
■24510▼aTrade-Offs and Opportunities in High-Dimensional Bayesian Modeling
■260 ▼a[Sl]▼bColumbia University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a259 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-02, Section: A.
■500 ▼aAdvisor: Gelman, Andrew;Rush, Cynthia.
■5021 ▼aThesis (Ph.D.)--Columbia University, 2024.
■520 ▼aWith the increasing availability of large multivariate datasets, modern parametric statistical models makes increasing use of high-dimensional parameter spaces to flexibly represent complex data generating mechanisms. Yet, ceteris paribus, increases in dimensionality often carry drawbacks across the various sub-problems of data analysis, posing challenges for the data analyst who must balance model plausibility against the practical considerations of implementation. We focus here on challenges to three components of data analysis: computation, inference, and model checking. In the computational domain, we are concerned with achieving reasonable scaling of the computational complexity with the parameter dimension without sacrificing the trustworthiness of our computation. Here, we study a particular class of algorithms - the vectorized approximate message passing (VAMP) iterations - which offer the possibility of linear per-iteration scaling with dimension. These iterations perform approximate inference for a class of Bayesian generalized linear regression models, and we demonstrate that under flexible distributional conditions, the estimation performance of these VAMP iterations can be predicted to high accuracy with probability decaying exponentially fast in the size of the regression problem. In the realm of statistical inference, we investigate the relationship between parameter dimension and identification. We develop formal notions of weak identification and model expansion in the Bayesian setting and use this to argue for a very general tendency for dimensionality-increasing model expansion to weaken the identification of model parameters. We draw two substantive conclusions from this formalism. First, the negative association between dimensionality and identification can be weakened or reversed when we construct prior distributions that encode sufficiently strong dependence between parameters. Absent such prior information, we derive bounds which indicate that decreasing identification is usually unavoidable with sufficient inflation of the dimension without increasing the severity of the third challenge we consider: that of dimensionality to model checking.We divide the topic of model checking into two sub-problems: fitness testing and correctness testing. Using our model expansion formalism, we show again that both of these problems tend to become more difficult as the model dimension grows. We propose two extensions of the posterior predictive \uD835\uDC5D-value - certain conditional and joint \uD835\uDC5D-values, which are designed to address these challenges for fitness and correctness testing respectively. We demonstrate the potential of these \uD835\uDC5D-values to allow successful model checking that scales with dimensionality theoretically and with examples.
■590 ▼aSchool code: 0054.
■650 4▼aStatistics
■650 4▼aStatistical physics
■650 4▼aInformation science
■653 ▼aBayesian statistics
■653 ▼aHigh-dimensional statistics
■653 ▼aInformation theory
■653 ▼aModel checking
■653 ▼aExpectation propagation algorithms
■690 ▼a0463
■690 ▼a0217
■690 ▼a0723
■71020▼aColumbia University▼bStatistics.
■7730 ▼tDissertations Abstracts International▼g86-02A.
■790 ▼a0054
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17162579▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


