Federated Learning (FL) allows clients to train a model collaboratively without sharing their private data. One key challenge in practical FL systems is data heterogeneity, particularly in handling clients with rare data, also referred to as Mavericks. These clients own one or more data classes exclusively, and the model performance becomes poor without their participation. Thus, utilizing Mavericks throughout training is crucial. In this paper, we first design a Maverick-aware Shapley valuation that fairly evaluates the contribution of Mavericks. The main idea is to compute the clients' Shapley values (SV) class-wise, i.e., per label. Next, we propose FedMS, a Maverick-Shapley client selection mechanism for FL that intelligently selects the clients that contribute the most in each round, by employing our Maverick-aware SV-based contribution score. We show that, compared to an extensive list of baselines, FedMS achieves better model performance and fairer Shapley Rewards distribution.
翻译:联邦学习(FL)允许客户端在不共享私有数据的情况下协作训练模型。实际FL系统中的关键挑战之一是数据异质性,尤其是在处理拥有罕见数据的客户端(也称为惊马)时。这些客户端独占一个或多个数据类别,若无其参与,模型性能将显著下降。因此,在整个训练过程中利用惊马至关重要。本文首先设计了一种惊马感知沙普利估值方法,用于公平评估惊马的贡献,其核心思想是按类别(即按标签)计算客户端的沙普利值(SV)。随后,我们提出FedMS——一种面向FL的惊马-沙普利客户端选择机制,通过采用基于惊马感知SV的贡献分数,智能选择每轮中贡献最大的客户端。我们证明,与大量基准方法相比,FedMS实现了更优的模型性能与更公平的沙普利奖励分配。