Federated Learning (FL) has emerged as a de facto machine learning area and received rapid increasing research interests from the community. However, catastrophic forgetting caused by data heterogeneity and partial participation poses distinctive challenges for FL, which are detrimental to the performance. To tackle the problems, we propose a new FL approach (namely GradMA), which takes inspiration from continual learning to simultaneously correct the server-side and worker-side update directions as well as take full advantage of server's rich computing and memory resources. Furthermore, we elaborate a memory reduction strategy to enable GradMA to accommodate FL with a large scale of workers. We then analyze convergence of GradMA theoretically under the smooth non-convex setting and show that its convergence rate achieves a linear speed up w.r.t the increasing number of sampled active workers. At last, our extensive experiments on various image classification tasks show that GradMA achieves significant performance gains in accuracy and communication efficiency compared to SOTA baselines.
翻译:联邦学习已成为机器学习领域的一种事实标准,并受到学界日益广泛的研究关注。然而,由数据异构性和部分参与导致的灾难性遗忘给联邦学习带来了独特挑战,对性能造成负面影响。为解决这些问题,我们提出了一种新的联邦学习方法(即GradMA),该方法从持续学习中汲取灵感,同时校正服务器端和工作器端的更新方向,并充分利用服务器丰富的计算和内存资源。此外,我们精心设计了一种内存缩减策略,使GradMA能够适应包含大量工作器的联邦学习场景。随后,我们在光滑非凸设置下对GradMA进行了理论收敛性分析,结果表明其收敛速率随采样活跃工作器数量的增加而实现线性加速。最后,我们在多种图像分类任务上的大量实验表明,与最先进基准方法相比,GradMA在准确率和通信效率方面均取得了显著性能提升。