Federated Learning (FL) has emerged as a de facto machine learning area and received rapid increasing research interests from the community. However, catastrophic forgetting caused by data heterogeneity and partial participation poses distinctive challenges for FL, which are detrimental to the performance. To tackle the problems, we propose a new FL approach (namely GradMA), which takes inspiration from continual learning to simultaneously correct the server-side and worker-side update directions as well as take full advantage of server's rich computing and memory resources. Furthermore, we elaborate a memory reduction strategy to enable GradMA to accommodate FL with a large scale of workers. We then analyze convergence of GradMA theoretically under the smooth non-convex setting and show that its convergence rate achieves a linear speed up w.r.t the increasing number of sampled active workers. At last, our extensive experiments on various image classification tasks show that GradMA achieves significant performance gains in accuracy and communication efficiency compared to SOTA baselines.
翻译:联邦学习已成为机器学习领域的重要范式,并引发了学术界日益增长的研究兴趣。然而,数据异构性和部分参与所引发的灾难性遗忘问题,为联邦学习带来了独特挑战,并严重损害其性能。为解决这些问题,本文提出一种新型联邦学习方法(命名为GradMA),该方法受持续学习启发,能够同时修正服务端与工作节点端的更新方向,并充分利用服务端丰富的计算与存储资源。此外,我们设计了一种内存缩减策略,使得GradMA能够适应大规模工作节点的联邦学习场景。随后,在光滑非凸设定下对GradMA进行理论收敛性分析,证明其收敛速率随采样的活跃工作节点数量增加而实现线性加速。最后,在多个图像分类任务上的大量实验表明,与当前最优基线方法相比,GradMA在准确率和通信效率方面均取得了显著性能提升。