서브메뉴
검색
Interpretable Algorithms for Data Integration: Adaptivity, Contamination, High-Dimensionality, and Privacy Constraints
Interpretable Algorithms for Data Integration: Adaptivity, Contamination, High-Dimensionality, and Privacy Constraints
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103552
- ISBN
- 9798280763029
- DDC
- 310
- 저자명
- Tian, Ye.
- 서명/저자
- Interpretable Algorithms for Data Integration: Adaptivity, Contamination, High-Dimensionality, and Privacy Constraints
- 발행사항
- [Sl] : Columbia University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 371 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-12, Section: A.
- 주기사항
- Advisor: Ying, Zhiliang.
- 학위논문주기
- Thesis (Ph.D.)--Columbia University, 2025.
- 초록/해제
- 요약This dissertation presents several methodological and theoretical contributions addressing modern challenges in data integration, with a focus on interpretable statistical approaches for tackling these issues.We begin by introducing the motivation for studying data integration, along with illustrative examples. We then outline key challenges encountered by existing methods, including distributional shifts, data contamination and adversarial (Byzantine) attacks, high-dimensional settings, and privacy constraints. The subsequent chapters address these challenges by introducing new methods with provable theoretical guarantees across various contexts.Our first contribution is a transfer learning algorithm tailored for high-dimensional generalized linear models, incorporating an aggregation step followed by a debiasing step. Building on this, we develop an inference procedure based on the debiased Lasso. We establish both finite-sample guarantees and asymptotic normality. In addition, we propose a transferable source detection method designed to identify and remove contaminated or anomalous datasets.Next, we propose a federated Gradient-EM algorithm for parameter estimation in general mixture models. The method is designed to be privacy-preserving, computationally efficient, and adaptive to model similarity. We show that, under certain conditions, the algorithm achieves nearly minimax-optimal estimation error rates within polynomial time. Our theoretical results also extend to mis-clustering error bounds in specific mixture model settings, and offer insight into the empirical success of related EM algorithms proposed in the literature.Finally, we introduce a representation learning framework for multi-task learning. Unlike prior work, our setting allows each task to have a distinct linear representation. We develop two algorithms for fitting this model via data integration, accommodating the presence of contaminated sources. One method is based on empirical risk minimization with a novel maximum principal angle penalty; the other leverages spectral methods. We derive both upper and lower bounds in finite samples, demonstrating near-optimal performance. The proposed framework generalizes prior models of linear representation and contributes several new theoretical and methodological insights.
- 일반주제명
- Statistics
- 일반주제명
- Information technology
- 키워드
- Data integration
- 키워드
- High dimension
- 키워드
- Robustness
- 기타저자
- Columbia University Statistics
- 기본자료저록
- Dissertations Abstracts International. 86-12A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017357729
■00520260202103552
■006m o d
■007cr#unu||||||||
■020 ▼a9798280763029
■035 ▼a(MiAaPQ)AAI32041895
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a310
■1001 ▼aTian, Ye.
■24510▼aInterpretable Algorithms for Data Integration: Adaptivity, Contamination, High-Dimensionality, and Privacy Constraints
■260 ▼a[Sl]▼bColumbia University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a371 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-12, Section: A.
■500 ▼aAdvisor: Ying, Zhiliang.
■5021 ▼aThesis (Ph.D.)--Columbia University, 2025.
■520 ▼aThis dissertation presents several methodological and theoretical contributions addressing modern challenges in data integration, with a focus on interpretable statistical approaches for tackling these issues.We begin by introducing the motivation for studying data integration, along with illustrative examples. We then outline key challenges encountered by existing methods, including distributional shifts, data contamination and adversarial (Byzantine) attacks, high-dimensional settings, and privacy constraints. The subsequent chapters address these challenges by introducing new methods with provable theoretical guarantees across various contexts.Our first contribution is a transfer learning algorithm tailored for high-dimensional generalized linear models, incorporating an aggregation step followed by a debiasing step. Building on this, we develop an inference procedure based on the debiased Lasso. We establish both finite-sample guarantees and asymptotic normality. In addition, we propose a transferable source detection method designed to identify and remove contaminated or anomalous datasets.Next, we propose a federated Gradient-EM algorithm for parameter estimation in general mixture models. The method is designed to be privacy-preserving, computationally efficient, and adaptive to model similarity. We show that, under certain conditions, the algorithm achieves nearly minimax-optimal estimation error rates within polynomial time. Our theoretical results also extend to mis-clustering error bounds in specific mixture model settings, and offer insight into the empirical success of related EM algorithms proposed in the literature.Finally, we introduce a representation learning framework for multi-task learning. Unlike prior work, our setting allows each task to have a distinct linear representation. We develop two algorithms for fitting this model via data integration, accommodating the presence of contaminated sources. One method is based on empirical risk minimization with a novel maximum principal angle penalty; the other leverages spectral methods. We derive both upper and lower bounds in finite samples, demonstrating near-optimal performance. The proposed framework generalizes prior models of linear representation and contributes several new theoretical and methodological insights.
■590 ▼aSchool code: 0054.
■650 4▼aStatistics
■650 4▼aInformation technology
■653 ▼aData integration
■653 ▼aFederated learning
■653 ▼aHigh dimension
■653 ▼aMinimax optimality
■653 ▼aRobustness
■653 ▼aTransfer learning
■690 ▼a0463
■690 ▼a0489
■690 ▼a0454
■690 ▼a0800
■71020▼aColumbia University▼bStatistics.
■7730 ▼tDissertations Abstracts International▼g86-12A.
■790 ▼a0054
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357729▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


