This paper presents the design and implementation of FLIPS, a middleware system to manage data and participant heterogeneity in federated learning (FL) training workloads. In particular, we examine the benefits of label distribution clustering on participant selection in federated learning. FLIPS clusters parties involved in an FL training job based on the label distribution of their data apriori, and during FL training, ensures that each cluster is equitably represented in the participants selected. FLIPS can support the most common FL algorithms, including FedAvg, FedProx, FedDyn, FedOpt and FedYogi. To manage platform heterogeneity and dynamic resource availability, FLIPS incorporates a straggler management mechanism to handle changing capacities in distributed, smart community applications. Privacy of label distributions, clustering and participant selection is ensured through a trusted execution environment (TEE). Our comprehensive empirical evaluation compares FLIPS with random participant selection, as well as three other "smart" selection mechanisms - Oort, TiFL and gradient clustering using two real-world datasets, two benchmark datasets, two different non-IID distributions and three common FL algorithms (FedYogi, FedProx and FedAvg). We demonstrate that FLIPS significantly improves convergence, achieving higher accuracy by 17 - 20 % with 20 - 60 % lower communication costs, and these benefits endure in the presence of straggler participants.
翻译:本文介绍了FLIPS的设计与实现,这是一个用于管理联邦学习训练工作负载中数据与参与者异构性的中间件系统。我们重点研究了标签分布聚类对联邦学习参与者选择的益处。FLIPS根据各方数据先验的标签分布对参与联邦学习训练任务的各方进行聚类,并在训练过程中确保每个聚类在所选参与者中具有公平的代表性。FLIPS支持最常用的联邦学习算法,包括FedAvg、FedProx、FedDyn、FedOpt和FedYogi。为管理平台异构性和动态资源可用性,FLIPS集成了掉队者管理机制,以应对分布式智能社区应用中变化的能力。通过可信执行环境保障标签分布、聚类和参与者选择的隐私性。我们全面的实证评估将FLIPS与随机参与者选择及其他三种"智能"选择机制(Oort、TiFL和梯度聚类)进行对比,使用了两个真实世界数据集、两个基准数据集、两种不同的非独立同分布场景及三种常见联邦学习算法(FedYogi、FedProx和FedAvg)。实验证明,FLIPS显著提升收敛性能,在降低20%-60%通信成本的同时实现17%-20%的精度提升,且这些优势在存在掉队者参与者时依然保持。