본문

서브메뉴

Advances in Probabilistic Machine Learning: Scalable Inference, Conditional Generation, and Invariance Modeling
Advances in Probabilistic Machine Learning: Scalable Inference, Conditional Generation, an...
Advances in Probabilistic Machine Learning: Scalable Inference, Conditional Generation, and Invariance Modeling

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202105201
ISBN  
9798297617254
DDC  
310
저자명  
Wu, Luhuan.
서명/저자  
Advances in Probabilistic Machine Learning: Scalable Inference, Conditional Generation, and Invariance Modeling
발행사항  
[Sl] : Columbia University, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
270 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-04, Section: B.
주기사항  
Advisor: Cunningham, John P.;Blei, David M.
학위논문주기  
Thesis (Ph.D.)--Columbia University, 2025.
초록/해제  
요약A central goal of machine learning is to uncover hidden patterns in the data for making predictions and drawing insights. The probabilistic perspective accounts for uncertainty by inferring a distribution over plausible patterns, while incorporating prior beliefs. However, applying probabilistic machine learning in modern settings presents several challenges, including scalability in large-data regimes, conditional generation with complex priors, and invariance modeling of heterogeneous data. This thesis develops methodologies to address these challenges. The first part of the thesis focuses on improving the scalability of Gaussian processes (GPs), a classical probabilistic model whose exact inference is intractable for large-scale problems. We first propose two approximate inference methods, one utilizing structured inducing points and the other exploiting sparsity in the prior precision matrix. While these methods are computationally attractive, they introduce biases that can affect downstream performance. In a separate line of work, we investigate systematic biases of two widely used scalable GP techniques and propose randomized algorithms to achieve unbiased inference. The second part of the thesis addresses inference challenges arising in modern deep generative models, in particular, diffusion models. These models capture distributions over complex data modalities, making them suitable as powerful priors for conditional generation tasks. However, inference from their conditional distributions is intractable. While previous methods rely on expensive training or error-prone approximations, we introduce a training-free sequential Monte Carlo algorithm that is asymptotically exact in the limit of increasing compute budget. We demonstrate the effectiveness of our algorithm on image generation and protein design applications. The third part of the thesis considers modeling challenges where data are collected from different environments. Fitting a model to pooled data may result in spurious correlations that fail to generalize to new environments. Instead, we aim to identify stable predictive relationships based on a subset of invariant features. To this end, we develop a probabilistic model for inferring invariant features with accompanying theoretical guarantees. To handle high-dimensional problems, we propose a scalable variational inference algorithm. Simulations and real-world experiments demonstrate improved inference accuracy and scalability over existing methods.
일반주제명  
Statistics
키워드  
Approximate inference
키워드  
Diffusion models
키워드  
Gaussian processes
키워드  
Invariant prediction
키워드  
Probabilistic machine learning
기타저자  
Columbia University Statistics
기본자료저록  
Dissertations Abstracts International. 87-04B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017359707
■00520260202105201
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798297617254
■035    ▼a(MiAaPQ)AAI32245053
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a310
■1001  ▼aWu,  Luhuan.
■24510▼aAdvances  in  Probabilistic  Machine  Learning:  Scalable  Inference,  Conditional  Generation,  and  Invariance  Modeling
■260    ▼a[Sl]▼bColumbia  University▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a270  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-04,  Section:  B.
■500    ▼aAdvisor:  Cunningham,  John  P.;Blei,  David  M.
■5021  ▼aThesis  (Ph.D.)--Columbia  University,  2025.
■520    ▼aA  central  goal  of  machine  learning  is  to  uncover  hidden  patterns  in  the  data  for  making  predictions  and  drawing  insights.  The  probabilistic  perspective  accounts  for  uncertainty  by  inferring  a  distribution  over  plausible  patterns,  while  incorporating  prior  beliefs.  However,  applying  probabilistic  machine  learning  in  modern  settings  presents  several  challenges,  including  scalability  in  large-data  regimes,  conditional  generation  with  complex  priors,  and  invariance  modeling  of  heterogeneous  data.  This  thesis  develops  methodologies  to  address  these  challenges.  The  first  part  of  the  thesis  focuses  on  improving  the  scalability  of  Gaussian  processes  (GPs),  a  classical  probabilistic  model  whose  exact  inference  is  intractable  for  large-scale  problems.  We  first  propose  two  approximate  inference  methods,  one  utilizing  structured  inducing  points  and  the  other  exploiting  sparsity  in  the  prior  precision  matrix.  While    these  methods  are  computationally  attractive,  they  introduce  biases  that  can  affect  downstream  performance.  In  a  separate  line  of  work,  we  investigate  systematic  biases  of  two  widely  used  scalable  GP  techniques  and  propose  randomized  algorithms  to  achieve  unbiased  inference.  The  second  part  of  the  thesis  addresses  inference  challenges  arising  in  modern  deep  generative  models,  in  particular,  diffusion  models.  These  models  capture  distributions  over  complex  data  modalities,    making  them  suitable  as  powerful  priors  for  conditional  generation  tasks.    However,  inference  from  their  conditional  distributions  is  intractable.  While  previous  methods  rely  on  expensive  training  or  error-prone    approximations,  we  introduce  a  training-free  sequential  Monte  Carlo  algorithm  that  is  asymptotically  exact  in  the  limit  of  increasing  compute  budget.  We  demonstrate  the  effectiveness  of  our  algorithm  on  image  generation  and  protein  design  applications.  The  third  part  of  the  thesis  considers  modeling  challenges  where  data  are  collected  from  different  environments.  Fitting  a  model  to  pooled  data  may  result  in  spurious  correlations  that  fail  to  generalize  to  new  environments.  Instead,  we  aim  to  identify  stable  predictive  relationships  based  on  a  subset  of  invariant  features.  To  this  end,  we  develop  a  probabilistic  model  for  inferring  invariant  features  with  accompanying  theoretical  guarantees.  To  handle  high-dimensional  problems,  we  propose  a  scalable  variational  inference  algorithm.  Simulations  and  real-world  experiments  demonstrate  improved  inference  accuracy  and  scalability  over  existing  methods.
■590    ▼aSchool  code:  0054.
■650  4▼aStatistics
■653    ▼aApproximate  inference
■653    ▼aDiffusion  models
■653    ▼aGaussian  processes
■653    ▼aInvariant  prediction
■653    ▼aProbabilistic  machine  learning
■690    ▼a0463
■690    ▼a0800
■71020▼aColumbia  University▼bStatistics.
■7730  ▼tDissertations  Abstracts  International▼g87-04B.
■790    ▼a0054
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17359707▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF14762 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.