In this paper, we investigate the problem of distributed learning (DL) in the presence of Byzantine attacks. For this problem, various robust bounded aggregation (RBA) rules have been proposed at the central server to mitigate the impact of Byzantine attacks. However, current DL methods apply RBA rules for the local gradients from the honest devices and the disruptive information from Byzantine devices, and the learning performance degrades significantly when the local gradients of different devices vary considerably from each other. To overcome this limitation, we propose a new DL method to cope with Byzantine attacks based on coded robust aggregation (CRA-DL). Before training begins, the training data are allocated to the devices redundantly. During training, in each iteration, the honest devices transmit coded gradients to the server computed from the allocated training data, and the server then aggregates the information received from both honest and Byzantine devices using RBA rules. In this way, the global gradient can be approximately recovered at the server to update the global model. Compared with current DL methods applying RBA rules, the improvement of CRA-DL is attributed to the fact that the coded gradients sent by the honest devices are closer to each other. This closeness enhances the robustness of the aggregation against Byzantine attacks, since Byzantine messages tend to be significantly different from those of honest devices in this case. We theoretically analyze the convergence performance of CRA-DL. Finally, we present numerical results to verify the superiority of the proposed method over existing baselines, showing its enhanced learning performance under Byzantine attacks.
翻译:本文研究了存在拜占庭攻击情况下的分布式学习问题。针对该问题,已有研究在中央服务器端提出了多种鲁棒有界聚合规则以减轻拜占庭攻击的影响。然而,现有分布式学习方法对来自诚实设备的本地梯度与来自拜占庭设备的破坏性信息均采用鲁棒有界聚合规则进行处理,当不同设备的本地梯度存在显著差异时,学习性能会严重下降。为克服这一局限,我们提出了一种基于编码鲁棒聚合的新型分布式学习方法以应对拜占庭攻击。在训练开始前,训练数据以冗余方式分配至各设备。在训练过程中,每个迭代周期内,诚实设备基于分配的训练数据计算编码梯度并传输至服务器,服务器随后使用鲁棒有界聚合规则对来自诚实设备与拜占庭设备的信息进行聚合。通过这种方式,服务器端可近似恢复全局梯度以更新全局模型。相较于当前应用鲁棒有界聚合规则的分布式学习方法,本方法的改进源于诚实设备发送的编码梯度彼此更为接近。这种接近性增强了聚合过程对拜占庭攻击的鲁棒性,因为在此情况下拜占庭消息往往与诚实设备消息存在显著差异。我们从理论上分析了该方法的收敛性能。最后,我们通过数值实验结果验证了所提方法相对于现有基线的优越性,展示了其在拜占庭攻击下增强的学习性能。