Decentralized learning (DL) has recently employed local updates to reduce the communication cost for general non-convex optimization problems. Specifically, local updates require each node to perform multiple update steps on the parameters of the local model before communicating with others. However, most existing methods could be highly sensitive to data heterogeneity (i.e., non-iid data distribution) and adversely affected by the stochastic gradient noise. In this paper, we propose DSE-MVR to address these problems.Specifically, DSE-MVR introduces a dual-slow estimation strategy that utilizes the gradient tracking technique to estimate the global accumulated update direction for handling the data heterogeneity problem; also for stochastic noise, the method uses the mini-batch momentum-based variance-reduction technique.We theoretically prove that DSE-MVR can achieve optimal convergence results for general non-convex optimization in both iid and non-iid data distribution settings. In particular, the leading terms in the convergence rates derived by DSE-MVR are independent of the stochastic noise for large-batches or large partial average intervals (i.e., the number of local update steps). Further, we put forward DSE-SGD and theoretically justify the importance of the dual-slow estimation strategy in the data heterogeneity setting. Finally, we conduct extensive experiments to show the superiority of DSE-MVR against other state-of-the-art approaches.
翻译:分散式学习(DL)近期采用局部更新来降低通用非凸优化问题的通信成本。具体而言,局部更新要求每个节点在与他人通信之前,对本地模型参数执行多次更新步骤。然而,现有大多数方法可能对数据异质性(即非独立同分布数据分布)高度敏感,并受随机梯度噪声的不利影响。本文提出DSE-MVR以解决上述问题。具体地,DSE-MVR引入双慢估计策略,利用梯度追踪技术估计全局累积更新方向以处理数据异质性;同时针对随机噪声,该方法采用基于小批量动量的方差缩减技术。我们从理论上证明,在独立同分布与非独立同分布数据分布设定下,DSE-MVR均能实现通用非凸优化的最优收敛结果。值得注意的是,DSE-MVR推导的收敛速率主导项在大批量或大部分平均间隔(即局部更新步数)条件下与随机噪声无关。此外,我们提出DSE-SGD并从理论上论证双慢估计策略在数据异质性场景中的重要性。最后,通过大量实验表明DSE-MVR相较于其他先进方法的优越性。