본문

서브메뉴

Interpretable Algorithms for Data Integration: Adaptivity, Contamination, High-Dimensionality, and Privacy Constraints
Interpretable Algorithms for Data Integration: Adaptivity, Contamination, High-Dimensional...
Interpretable Algorithms for Data Integration: Adaptivity, Contamination, High-Dimensionality, and Privacy Constraints

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202103552
ISBN  
9798280763029
DDC  
310
저자명  
Tian, Ye.
서명/저자  
Interpretable Algorithms for Data Integration: Adaptivity, Contamination, High-Dimensionality, and Privacy Constraints
발행사항  
[Sl] : Columbia University, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
371 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-12, Section: A.
주기사항  
Advisor: Ying, Zhiliang.
학위논문주기  
Thesis (Ph.D.)--Columbia University, 2025.
초록/해제  
요약This dissertation presents several methodological and theoretical contributions addressing modern challenges in data integration, with a focus on interpretable statistical approaches for tackling these issues.We begin by introducing the motivation for studying data integration, along with illustrative examples. We then outline key challenges encountered by existing methods, including distributional shifts, data contamination and adversarial (Byzantine) attacks, high-dimensional settings, and privacy constraints. The subsequent chapters address these challenges by introducing new methods with provable theoretical guarantees across various contexts.Our first contribution is a transfer learning algorithm tailored for high-dimensional generalized linear models, incorporating an aggregation step followed by a debiasing step. Building on this, we develop an inference procedure based on the debiased Lasso. We establish both finite-sample guarantees and asymptotic normality. In addition, we propose a transferable source detection method designed to identify and remove contaminated or anomalous datasets.Next, we propose a federated Gradient-EM algorithm for parameter estimation in general mixture models. The method is designed to be privacy-preserving, computationally efficient, and adaptive to model similarity. We show that, under certain conditions, the algorithm achieves nearly minimax-optimal estimation error rates within polynomial time. Our theoretical results also extend to mis-clustering error bounds in specific mixture model settings, and offer insight into the empirical success of related EM algorithms proposed in the literature.Finally, we introduce a representation learning framework for multi-task learning. Unlike prior work, our setting allows each task to have a distinct linear representation. We develop two algorithms for fitting this model via data integration, accommodating the presence of contaminated sources. One method is based on empirical risk minimization with a novel maximum principal angle penalty; the other leverages spectral methods. We derive both upper and lower bounds in finite samples, demonstrating near-optimal performance. The proposed framework generalizes prior models of linear representation and contributes several new theoretical and methodological insights.
일반주제명  
Statistics
일반주제명  
Information technology
키워드  
Data integration
키워드  
Federated learning
키워드  
High dimension
키워드  
Minimax optimality
키워드  
Robustness
키워드  
Transfer learning
기타저자  
Columbia University Statistics
기본자료저록  
Dissertations Abstracts International. 86-12A.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017357729
■00520260202103552
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798280763029
■035    ▼a(MiAaPQ)AAI32041895
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a310
■1001  ▼aTian,  Ye.
■24510▼aInterpretable  Algorithms  for  Data  Integration:  Adaptivity,  Contamination,  High-Dimensionality,  and  Privacy  Constraints
■260    ▼a[Sl]▼bColumbia  University▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a371  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-12,  Section:  A.
■500    ▼aAdvisor:  Ying,  Zhiliang.
■5021  ▼aThesis  (Ph.D.)--Columbia  University,  2025.
■520    ▼aThis  dissertation  presents  several  methodological  and  theoretical  contributions  addressing  modern  challenges  in  data  integration,  with  a  focus  on  interpretable  statistical  approaches  for  tackling  these  issues.We  begin  by  introducing  the  motivation  for  studying  data  integration,  along  with  illustrative  examples.  We  then  outline  key  challenges  encountered  by  existing  methods,  including  distributional  shifts,  data  contamination  and  adversarial  (Byzantine)  attacks,  high-dimensional  settings,  and  privacy  constraints.  The  subsequent  chapters  address  these  challenges  by  introducing  new  methods  with  provable  theoretical  guarantees  across  various  contexts.Our  first  contribution  is  a  transfer  learning  algorithm  tailored  for  high-dimensional  generalized  linear  models,  incorporating  an  aggregation  step  followed  by  a  debiasing  step.  Building  on  this,  we  develop  an  inference  procedure  based  on  the  debiased  Lasso.  We  establish  both  finite-sample  guarantees  and  asymptotic  normality.  In  addition,  we  propose  a  transferable  source  detection  method  designed  to  identify  and  remove  contaminated  or  anomalous  datasets.Next,  we  propose  a  federated  Gradient-EM  algorithm  for  parameter  estimation  in  general  mixture  models.  The  method  is  designed  to  be  privacy-preserving,  computationally  efficient,  and  adaptive  to  model  similarity.  We  show  that,  under  certain  conditions,  the  algorithm  achieves  nearly  minimax-optimal  estimation  error  rates  within  polynomial  time.  Our  theoretical  results  also  extend  to  mis-clustering  error  bounds  in  specific  mixture  model  settings,  and  offer  insight  into  the  empirical  success  of  related  EM  algorithms  proposed  in  the  literature.Finally,  we  introduce  a  representation  learning  framework  for  multi-task  learning.  Unlike  prior  work,  our  setting  allows  each  task  to  have  a  distinct  linear  representation.  We  develop  two  algorithms  for  fitting  this  model  via  data  integration,  accommodating  the  presence  of  contaminated  sources.  One  method  is  based  on  empirical  risk  minimization  with  a  novel  maximum  principal  angle  penalty;  the  other  leverages  spectral  methods.  We  derive  both  upper  and  lower  bounds  in  finite  samples,  demonstrating  near-optimal  performance.  The  proposed  framework  generalizes  prior  models  of  linear  representation  and  contributes  several  new  theoretical  and  methodological  insights.
■590    ▼aSchool  code:  0054.
■650  4▼aStatistics
■650  4▼aInformation  technology
■653    ▼aData  integration
■653    ▼aFederated  learning
■653    ▼aHigh  dimension
■653    ▼aMinimax  optimality
■653    ▼aRobustness
■653    ▼aTransfer  learning
■690    ▼a0463
■690    ▼a0489
■690    ▼a0454
■690    ▼a0800
■71020▼aColumbia  University▼bStatistics.
■7730  ▼tDissertations  Abstracts  International▼g86-12A.
■790    ▼a0054
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357729▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF18958 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.