Training modern neural networks or models typically requires averaging over a sample of high-dimensional vectors. Poisoning attacks can skew or bias the average vectors used to train the model, forcing the model to learn specific patterns or avoid learning anything useful. Byzantine robust aggregation is a principled algorithmic defense against such biasing. Robust aggregators can bound the maximum bias in computing centrality statistics, such as mean, even when some fraction of inputs are arbitrarily corrupted. Designing such aggregators is challenging when dealing with high dimensions. However, the first polynomial-time algorithms with strong theoretical bounds on the bias have recently been proposed. Their bounds are independent of the number of dimensions, promising a conceptual limit on the power of poisoning attacks in their ongoing arms race against defenses. In this paper, we show a new attack called HIDRA on practical realization of strong defenses which subverts their claim of dimension-independent bias. HIDRA highlights a novel computational bottleneck that has not been a concern of prior information-theoretic analysis. Our experimental evaluation shows that our attacks almost completely destroy the model performance, whereas existing attacks with the same goal fail to have much effect. Our findings leave the arms race between poisoning attacks and provable defenses wide open.
翻译:现代神经网络或模型的训练通常需要对高维向量样本进行平均。投毒攻击能够扭曲或偏置用于训练模型的平均向量,迫使模型学习特定模式或无法学习有效内容。拜占庭鲁棒聚合是针对此类偏置的一种原则性算法防御手段。即使部分输入被任意篡改,鲁棒聚合器仍能限制均值等中心性统计量的最大偏差。设计此类聚合器在高维场景中极具挑战性。然而,近期首次提出了具有强理论偏差界且时间复杂度为多项式的算法,其偏差界与维度无关,这为投毒攻击与防御之间持续军备竞赛中的攻击能力设定了概念性极限。本文提出一种名为HIDRA的新攻击方法,针对强防御的实际实现,颠覆了其声称的维度无关偏差特性。HIDRA揭示了前人信息论分析未曾关注的新的计算瓶颈。实验评估表明,我们的攻击几乎完全摧毁了模型性能,而具有相同目标的现有攻击却收效甚微。这一发现使得投毒攻击与可证明防御之间的军备竞赛仍充满变数。