Federated learning is a decentralized machine learning framework wherein not all clients are able to participate in each round. An emerging line of research is devoted to tackling arbitrary client unavailability. Existing theoretical analysis imposes restrictive structural assumptions on the unavailability patterns, and their proposed algorithms were tailored to those assumptions. In this paper, we relax those assumptions and consider adversarial client unavailability. To quantify the degrees of client unavailability, we use the notion of {\em $\epsilon$-adversary dropout fraction}. For both non-convex and strongly-convex global objectives, we show that simple variants of FedAvg or FedProx, albeit completely agnostic to $\epsilon$, converge to an estimation error on the order of $\epsilon (G^2 + \sigma^2)$, where $G$ is a heterogeneity parameter and $\sigma^2$ is the noise level. We prove that this estimation error is minimax-optimal. We also show that the variants of FedAvg or FedProx have convergence speeds $O(1/\sqrt{T})$ for non-convex objectives and $O(1/T)$ for strongly-convex objectives, both of which are the best possible for any first-order method that only has access to noisy gradients. Our proofs build upon a tight analysis of the selection bias that persists in the entire learning process. We validate our theoretical prediction through numerical experiments on synthetic and real-world datasets.
翻译:联邦学习是一种去中心化的机器学习框架,其中并非所有客户端都能参与每一轮训练。新兴的研究方向致力于解决任意客户端不可用问题。现有理论分析对不可用模式施加了严格的结构性假设,其提出的算法也针对这些假设进行了定制。本文放宽了这些假设,并考虑敌对客户端不可用情况。为量化客户端不可用程度,我们采用{\em $\epsilon$-敌对丢弃比例}这一概念。对于非凸和强凸全局目标函数,我们证明FedAvg或FedProx的简单变体——尽管完全不知道$\epsilon$——收敛到阶为$\epsilon (G^2 + \sigma^2)$的估计误差,其中$G$为异质性参数,$\sigma^2$为噪声水平。我们证明该估计误差是最小最大最优的。我们还表明,FedAvg或FedProx的变体对于非凸目标函数具有$O(1/\sqrt{T})$的收敛速度,对于强凸目标函数具有$O(1/T)$的收敛速度,两者均为仅能访问噪声梯度的一阶方法所能达到的最优速度。我们的证明建立在对整个学习过程中持续存在的选择偏差的严格分析之上。通过合成数据集和真实数据集的数值实验,我们验证了理论预测。