We present a data poisoning attack -- Phantom Transfer -- with the property that, even if you know precisely how the poison was placed into an otherwise benign dataset, you cannot filter it out. We achieve this by modifying subliminal learning to work in real-world contexts and demonstrate that the attack works regardless of which model produced the data, which model is trained on the data or what the attack target is. Furthermore, the attack survives 11 tested data-level defences, including one where every sample is paraphrased by another model. We characterise when this attack works best and show that it can be used to plant password-triggered behaviours into models while still beating defences. In short, we provide an existence proof that maximum-affordance defences can fail to stop sophisticated data poisoning attacks. We suggest that future defences should be supplemented with white-box methods and post-training model audits.
翻译:我们提出一种名为“幻影传递”的数据投毒攻击方法,其特性在于:即便防御方精确知晓有毒样本如何被植入原本良性的数据集中,仍无法将其过滤剔除。通过将阈下学习机制改造为适用于现实场景,我们证明了该攻击的有效性不受以下因素影响:生成数据的模型、基于该数据训练的模型类型乃至攻击目标。此外,该攻击能突破11种经测试的数据级防御措施——包括使用另一模型对每个样本进行语义复述的重写防御。我们刻画了该攻击的最佳生效条件,并展示其可在突破防御的同时,将密码触发行为植入模型。简言之,本工作提供了存在性证明:最大权限防御机制亦无法阻断高级数据投毒攻击。我们建议未来的防御应辅以白盒方法与训练后模型审计。