서브메뉴
검색
Advances in Probabilistic Machine Learning: Scalable Inference, Conditional Generation, and Invariance Modeling
Advances in Probabilistic Machine Learning: Scalable Inference, Conditional Generation, and Invariance Modeling
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105201
- ISBN
- 9798297617254
- DDC
- 310
- 저자명
- Wu, Luhuan.
- 서명/저자
- Advances in Probabilistic Machine Learning: Scalable Inference, Conditional Generation, and Invariance Modeling
- 발행사항
- [Sl] : Columbia University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 270 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-04, Section: B.
- 주기사항
- Advisor: Cunningham, John P.;Blei, David M.
- 학위논문주기
- Thesis (Ph.D.)--Columbia University, 2025.
- 초록/해제
- 요약A central goal of machine learning is to uncover hidden patterns in the data for making predictions and drawing insights. The probabilistic perspective accounts for uncertainty by inferring a distribution over plausible patterns, while incorporating prior beliefs. However, applying probabilistic machine learning in modern settings presents several challenges, including scalability in large-data regimes, conditional generation with complex priors, and invariance modeling of heterogeneous data. This thesis develops methodologies to address these challenges. The first part of the thesis focuses on improving the scalability of Gaussian processes (GPs), a classical probabilistic model whose exact inference is intractable for large-scale problems. We first propose two approximate inference methods, one utilizing structured inducing points and the other exploiting sparsity in the prior precision matrix. While these methods are computationally attractive, they introduce biases that can affect downstream performance. In a separate line of work, we investigate systematic biases of two widely used scalable GP techniques and propose randomized algorithms to achieve unbiased inference. The second part of the thesis addresses inference challenges arising in modern deep generative models, in particular, diffusion models. These models capture distributions over complex data modalities, making them suitable as powerful priors for conditional generation tasks. However, inference from their conditional distributions is intractable. While previous methods rely on expensive training or error-prone approximations, we introduce a training-free sequential Monte Carlo algorithm that is asymptotically exact in the limit of increasing compute budget. We demonstrate the effectiveness of our algorithm on image generation and protein design applications. The third part of the thesis considers modeling challenges where data are collected from different environments. Fitting a model to pooled data may result in spurious correlations that fail to generalize to new environments. Instead, we aim to identify stable predictive relationships based on a subset of invariant features. To this end, we develop a probabilistic model for inferring invariant features with accompanying theoretical guarantees. To handle high-dimensional problems, we propose a scalable variational inference algorithm. Simulations and real-world experiments demonstrate improved inference accuracy and scalability over existing methods.
- 일반주제명
- Statistics
- 키워드
- Diffusion models
- 기타저자
- Columbia University Statistics
- 기본자료저록
- Dissertations Abstracts International. 87-04B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017359707
■00520260202105201
■006m o d
■007cr#unu||||||||
■020 ▼a9798297617254
■035 ▼a(MiAaPQ)AAI32245053
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a310
■1001 ▼aWu, Luhuan.
■24510▼aAdvances in Probabilistic Machine Learning: Scalable Inference, Conditional Generation, and Invariance Modeling
■260 ▼a[Sl]▼bColumbia University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a270 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-04, Section: B.
■500 ▼aAdvisor: Cunningham, John P.;Blei, David M.
■5021 ▼aThesis (Ph.D.)--Columbia University, 2025.
■520 ▼aA central goal of machine learning is to uncover hidden patterns in the data for making predictions and drawing insights. The probabilistic perspective accounts for uncertainty by inferring a distribution over plausible patterns, while incorporating prior beliefs. However, applying probabilistic machine learning in modern settings presents several challenges, including scalability in large-data regimes, conditional generation with complex priors, and invariance modeling of heterogeneous data. This thesis develops methodologies to address these challenges. The first part of the thesis focuses on improving the scalability of Gaussian processes (GPs), a classical probabilistic model whose exact inference is intractable for large-scale problems. We first propose two approximate inference methods, one utilizing structured inducing points and the other exploiting sparsity in the prior precision matrix. While these methods are computationally attractive, they introduce biases that can affect downstream performance. In a separate line of work, we investigate systematic biases of two widely used scalable GP techniques and propose randomized algorithms to achieve unbiased inference. The second part of the thesis addresses inference challenges arising in modern deep generative models, in particular, diffusion models. These models capture distributions over complex data modalities, making them suitable as powerful priors for conditional generation tasks. However, inference from their conditional distributions is intractable. While previous methods rely on expensive training or error-prone approximations, we introduce a training-free sequential Monte Carlo algorithm that is asymptotically exact in the limit of increasing compute budget. We demonstrate the effectiveness of our algorithm on image generation and protein design applications. The third part of the thesis considers modeling challenges where data are collected from different environments. Fitting a model to pooled data may result in spurious correlations that fail to generalize to new environments. Instead, we aim to identify stable predictive relationships based on a subset of invariant features. To this end, we develop a probabilistic model for inferring invariant features with accompanying theoretical guarantees. To handle high-dimensional problems, we propose a scalable variational inference algorithm. Simulations and real-world experiments demonstrate improved inference accuracy and scalability over existing methods.
■590 ▼aSchool code: 0054.
■650 4▼aStatistics
■653 ▼aApproximate inference
■653 ▼aDiffusion models
■653 ▼aGaussian processes
■653 ▼aInvariant prediction
■653 ▼aProbabilistic machine learning
■690 ▼a0463
■690 ▼a0800
■71020▼aColumbia University▼bStatistics.
■7730 ▼tDissertations Abstracts International▼g87-04B.
■790 ▼a0054
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17359707▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


