Distributed learning has emerged as a leading paradigm for training large machine learning models. However, in real-world scenarios, participants may be unreliable or malicious, posing a significant challenge to the integrity and accuracy of the trained models. Byzantine fault tolerance mechanisms have been proposed to address these issues, but they often assume full participation from all clients, which is not always practical due to the unavailability of some clients or communication constraints. In our work, we propose the first distributed method with client sampling and provable tolerance to Byzantine workers. The key idea behind the developed method is the use of gradient clipping to control stochastic gradient differences in recursive variance reduction. This allows us to bound the potential harm caused by Byzantine workers, even during iterations when all sampled clients are Byzantine. Furthermore, we incorporate communication compression into the method to enhance communication efficiency. Under general assumptions, we prove convergence rates for the proposed method that match the existing state-of-the-art (SOTA) theoretical results. We also propose a heuristic on adjusting any Byzantine-robust method to a partial participation scenario via clipping.
翻译:分布式学习已成为训练大型机器学习模型的主流范式。然而,在实际场景中,参与者可能不可靠或具有恶意性,这对训练模型的完整性与准确性构成了重大挑战。拜占庭容错机制已被提出以解决这些问题,但这些机制通常假设所有客户端完全参与,由于部分客户端不可用或通信限制,该假设在实际中往往难以成立。在本研究中,我们提出了首个支持客户端采样且可证明容忍拜占庭工作节点的分布式方法。该方法的核心思想是利用梯度裁剪技术控制递归方差缩减中的随机梯度差异。这使得我们能够限制拜占庭工作节点可能造成的损害,即使在所有采样客户端均为拜占庭节点的迭代轮次中亦然。此外,我们在方法中引入了通信压缩机制以提升通信效率。在一般性假设下,我们证明了所提方法的收敛速率与现有最先进(SOTA)理论结果相匹配。我们还提出了一种启发式策略,通过裁剪技术将任意拜占庭鲁棒方法适配至部分参与场景。