서브메뉴
검색
Personalized and Distributed Data Analytics in Heterogeneous Environments
Personalized and Distributed Data Analytics in Heterogeneous Environments
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103650
- ISBN
- 9798314875759
- DDC
- 658
- 저자명
- Shi, Naichen.
- 서명/저자
- Personalized and Distributed Data Analytics in Heterogeneous Environments
- 발행사항
- [Sl] : University of Michigan, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 248 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-11, Section: B.
- 주기사항
- Advisor: Al Kontar, Raed.
- 학위논문주기
- Thesis (Ph.D.)--University of Michigan, 2025.
- 초록/해제
- 요약It is a common wisdom in statistics that more data leads to better models. However, as data are increasingly collected from distributed sources, such as different devices or users, their inherent statistical heterogeneity creates challenges for effective knowledge integration. Conventional population-based models often rely on i.i.d. assumptions, which often neglect variations across data sources. When data distributions differ, understanding their structure and integrating information for predictive modeling becomes non-trivial. This dissertation tackles these challenges through personalized modeling. Instead of fitting one single model for data from all sources, personalized data analytics fits data source-specific models while still encouraging knowledge transfer across sources. This dissertation proposes personalized descriptive and predictive analytics that attempt to answer three key questions: (Q1) How can we develop descriptive analytics to extract shared and unique patterns from heterogeneous data? (Q2) How can we design robust statistical methods that remain reliable in the presence of outliers? (Q3) How can we leverage insights from covariate and concept shifts to construct effective personalized predictive models? To answer these questions, the dissertation proposes three methodological contributions. Chapter 2 proposes Personalized PCA (PerPCA), a novel approach that distinguishes shared and unique features across data sources using mutually orthogonal global and local principal components. Chapter 3 presents Triple Component Matrix Factorization (TCMF) to recover global, local, and noisy components in multi-source data corrupted by outlier noise. Both PerPCA and TCMF are equipped with theoretical guarantees on statistical errors. Chapter 4 develops a predictive modeling framework called Personalized Federated Learning via Domain Adaptation (PFL-DA) that addresses both covariate and concept shifts across distributed sources. The proposed methods provide scalable and interpretable solutions for extracting insights, integrating knowledge, and improving predictive performance in distributed and heterogeneous environments. These findings have broad applications across various domains, including image and video processing, topic modeling, and manufacturing.
- 일반주제명
- Industrial engineering
- 일반주제명
- Computer engineering
- 일반주제명
- Engineering
- 키워드
- Data analytics
- 기타저자
- University of Michigan Industrial & Operations Engineering
- 기본자료저록
- Dissertations Abstracts International. 86-11B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017358143
■00520260202103650
■006m o d
■007cr#unu||||||||
■020 ▼a9798314875759
■035 ▼a(MiAaPQ)AAI32092682
■035 ▼a(MiAaPQ)umichrackham006125
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a658
■1001 ▼aShi, Naichen.
■24510▼aPersonalized and Distributed Data Analytics in Heterogeneous Environments
■260 ▼a[Sl]▼bUniversity of Michigan▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a248 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-11, Section: B.
■500 ▼aAdvisor: Al Kontar, Raed.
■5021 ▼aThesis (Ph.D.)--University of Michigan, 2025.
■520 ▼aIt is a common wisdom in statistics that more data leads to better models. However, as data are increasingly collected from distributed sources, such as different devices or users, their inherent statistical heterogeneity creates challenges for effective knowledge integration. Conventional population-based models often rely on i.i.d. assumptions, which often neglect variations across data sources. When data distributions differ, understanding their structure and integrating information for predictive modeling becomes non-trivial. This dissertation tackles these challenges through personalized modeling. Instead of fitting one single model for data from all sources, personalized data analytics fits data source-specific models while still encouraging knowledge transfer across sources. This dissertation proposes personalized descriptive and predictive analytics that attempt to answer three key questions: (Q1) How can we develop descriptive analytics to extract shared and unique patterns from heterogeneous data? (Q2) How can we design robust statistical methods that remain reliable in the presence of outliers? (Q3) How can we leverage insights from covariate and concept shifts to construct effective personalized predictive models? To answer these questions, the dissertation proposes three methodological contributions. Chapter 2 proposes Personalized PCA (PerPCA), a novel approach that distinguishes shared and unique features across data sources using mutually orthogonal global and local principal components. Chapter 3 presents Triple Component Matrix Factorization (TCMF) to recover global, local, and noisy components in multi-source data corrupted by outlier noise. Both PerPCA and TCMF are equipped with theoretical guarantees on statistical errors. Chapter 4 develops a predictive modeling framework called Personalized Federated Learning via Domain Adaptation (PFL-DA) that addresses both covariate and concept shifts across distributed sources. The proposed methods provide scalable and interpretable solutions for extracting insights, integrating knowledge, and improving predictive performance in distributed and heterogeneous environments. These findings have broad applications across various domains, including image and video processing, topic modeling, and manufacturing.
■590 ▼aSchool code: 0127.
■650 4▼aIndustrial engineering
■650 4▼aComputer engineering
■650 4▼aEngineering
■653 ▼aPersonalized modeling
■653 ▼aData analytics
■653 ▼aFeature extraction
■653 ▼aTriple Component Matrix Factorization
■690 ▼a0546
■690 ▼a0796
■690 ▼a0464
■690 ▼a0537
■71020▼aUniversity of Michigan▼bIndustrial & Operations Engineering.
■7730 ▼tDissertations Abstracts International▼g86-11B.
■790 ▼a0127
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17358143▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


