서브메뉴
검색
Communication and Data Efficient Algorithms for Vertical Federated Learning
Communication and Data Efficient Algorithms for Vertical Federated Learning
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202104804
- ISBN
- 9798290938615
- DDC
- 004
- 저자명
- Valdeira, Pedro.
- 서명/저자
- Communication and Data Efficient Algorithms for Vertical Federated Learning
- 발행사항
- [Sl] : Carnegie Mellon University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 131 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-02, Section: A.
- 주기사항
- Advisor: Chi, Yuejie.
- 학위논문주기
- Thesis (Ph.D.)--Carnegie Mellon University, 2025.
- 초록/해제
- 요약Federated learning (FL) is a collaborative machine learning paradigm where multiple clients jointly train a model without sharing their raw local data. FL can be categorized as horizontal FL, where clients hold different subsets of samples, or vertical FL (VFL), where clients hold different features of a shared set of samples. There have been significant recent advances in FL, yet most research has focused on the horizontal setting. Despite the importance of VFL to domains such as finance and healthcare, there is still a gap between existing work and real-world setups-namely when it comes to dealing with communication and data-availability constraints.The first part of this thesis focuses on communication efficiency in VFL. In FL, client-server schemes, where the clients communicate only with a central server, are predominant; however, such setups strain the finite bandwidth of the server, often causing latency and becoming a bottleneck in training. To mitigate this, we resort to lossy compression. We introduce error feedback compressed VFL (EF-VFL), which employs an error feedback mechanism to stabilize communication-compressed training in VFL. Numerical experiments confirm the improved communication efficiency of EF-VFL over state-of-the-art methods.The second part of this thesis addresses data efficiency in VFL. More precisely, we tackle the problem of leveraging incomplete samples, where some clients lack their feature blocks, during both training and inference. We propose LASER-VFL, a simple yet effective method relying on model parameter sharing and a task-sampling mechanism to efficiently train a family of predictors that is capable of handling arbitrary sets of observed feature blocks. LASER-VFL significantly improves over the performance of baselines, both in the presence and, remarkably, even in the absence of missing features.The third part of this thesis explores an alternative approach to mitigate the bandwidth bottleneck in the client-server scheme: the use of direct client-client links. In this semi-decentralized scheme, client-client links enable information flow without straining server bandwidth, while client-server links accelerate convergence in poorly connected client networks. We propose multi-token coordinate descent (MTCD), a semi-decentralized VFL method. Token methods avoid the staleness seen in most decentralized optimization methods, at the cost of losing their parallelism. By running multiple tokens in parallel, MTCD allow us to navigate this staleness-parallelism trade-off. MTCD recovers the client-server and decentralized schemes as special cases, and, by tuning its reliance on client-server links, it can also span the spectrum of semi-decentralized configurations in between. This flexibility allows MTCD to outperform both ends of the spectrum in applications where neither extreme is ideal.
- 일반주제명
- Computer science
- 일반주제명
- Information science
- 키워드
- Missing features
- 기타저자
- Carnegie Mellon University Electrical and Computer Engineering
- 기본자료저록
- Dissertations Abstracts International. 87-02A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017358878
■00520260202104804
■006m o d
■007cr#unu||||||||
■020 ▼a9798290938615
■035 ▼a(MiAaPQ)AAI32165024
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aValdeira, Pedro.▼0(orcid)0000-0003-2677-8710
■24510▼aCommunication and Data Efficient Algorithms for Vertical Federated Learning
■260 ▼a[Sl]▼bCarnegie Mellon University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a131 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-02, Section: A.
■500 ▼aAdvisor: Chi, Yuejie.
■5021 ▼aThesis (Ph.D.)--Carnegie Mellon University, 2025.
■520 ▼aFederated learning (FL) is a collaborative machine learning paradigm where multiple clients jointly train a model without sharing their raw local data. FL can be categorized as horizontal FL, where clients hold different subsets of samples, or vertical FL (VFL), where clients hold different features of a shared set of samples. There have been significant recent advances in FL, yet most research has focused on the horizontal setting. Despite the importance of VFL to domains such as finance and healthcare, there is still a gap between existing work and real-world setups-namely when it comes to dealing with communication and data-availability constraints.The first part of this thesis focuses on communication efficiency in VFL. In FL, client-server schemes, where the clients communicate only with a central server, are predominant; however, such setups strain the finite bandwidth of the server, often causing latency and becoming a bottleneck in training. To mitigate this, we resort to lossy compression. We introduce error feedback compressed VFL (EF-VFL), which employs an error feedback mechanism to stabilize communication-compressed training in VFL. Numerical experiments confirm the improved communication efficiency of EF-VFL over state-of-the-art methods.The second part of this thesis addresses data efficiency in VFL. More precisely, we tackle the problem of leveraging incomplete samples, where some clients lack their feature blocks, during both training and inference. We propose LASER-VFL, a simple yet effective method relying on model parameter sharing and a task-sampling mechanism to efficiently train a family of predictors that is capable of handling arbitrary sets of observed feature blocks. LASER-VFL significantly improves over the performance of baselines, both in the presence and, remarkably, even in the absence of missing features.The third part of this thesis explores an alternative approach to mitigate the bandwidth bottleneck in the client-server scheme: the use of direct client-client links. In this semi-decentralized scheme, client-client links enable information flow without straining server bandwidth, while client-server links accelerate convergence in poorly connected client networks. We propose multi-token coordinate descent (MTCD), a semi-decentralized VFL method. Token methods avoid the staleness seen in most decentralized optimization methods, at the cost of losing their parallelism. By running multiple tokens in parallel, MTCD allow us to navigate this staleness-parallelism trade-off. MTCD recovers the client-server and decentralized schemes as special cases, and, by tuning its reliance on client-server links, it can also span the spectrum of semi-decentralized configurations in between. This flexibility allows MTCD to outperform both ends of the spectrum in applications where neither extreme is ideal.
■590 ▼aSchool code: 0041.
■650 4▼aComputer science
■650 4▼aInformation science
■653 ▼aCommunication-compressed optimization
■653 ▼aDistributed optimization
■653 ▼aMissing features
■653 ▼aNonconvex optimization
■653 ▼aVertical federated learning
■690 ▼a0984
■690 ▼a0800
■690 ▼a0723
■71020▼aCarnegie Mellon University▼bElectrical and Computer Engineering.
■7730 ▼tDissertations Abstracts International▼g87-02A.
■790 ▼a0041
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17358878▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


