This paper studies Byzantine-robust stochastic optimization over a decentralized network, where every agent periodically communicates with its neighbors to exchange local models, and then updates its own local model by stochastic gradient descent (SGD). The performance of such a method is affected by an unknown number of Byzantine agents, which conduct adversarially during the optimization process. To the best of our knowledge, there is no existing work that simultaneously achieves a linear convergence speed and a small learning error. We observe that the learning error is largely dependent on the intrinsic stochastic gradient noise. Motivated by this observation, we introduce two variance reduction methods, stochastic average gradient algorithm (SAGA) and loopless stochastic variance-reduced gradient (LSVRG), to Byzantine-robust decentralized stochastic optimization for eliminating the negative effect of the stochastic gradient noise. The two resulting methods, BRAVO-SAGA and BRAVO-LSVRG, enjoy both linear convergence speeds and stochastic gradient noise-independent learning errors. Such learning errors are optimal for a class of methods based on total variation (TV)-norm regularization and stochastic subgradient update. We conduct extensive numerical experiments to demonstrate their effectiveness under various Byzantine attacks.
翻译:本文研究去中心化网络上的拜占庭鲁棒随机优化问题,其中每个智能体定期与邻居通信以交换局部模型,随后通过随机梯度下降(SGD)更新自身局部模型。此类方法的性能受到未知数量的拜占庭智能体影响,这些智能体在优化过程中会进行对抗性操作。据我们所知,目前尚无工作能同时实现线性收敛速度与较小学习误差。我们观察到学习误差在很大程度上取决于内在的随机梯度噪声。基于此发现,我们引入两种方差缩减方法——随机平均梯度算法(SAGA)与无环随机方差缩减梯度(LSVRG)——用于拜占庭鲁棒去中心化随机优化,以消除随机梯度噪声的负面影响。由此产生的两种方法BRAVO-SAGA与BRAVO-LSVRG兼具线性收敛速度与随机梯度噪声无关的学习误差。对于基于全变差(TV)范数正则化与随机次梯度更新的一类方法而言,此类学习误差是最优的。我们进行了大量数值实验,以证明它们在多种拜占庭攻击下的有效性。