Targeted data poisoning attacks manipulate model predictions on specific test samples by injecting malicious data into training. Yet existing evaluations report average attack success rates over randomly selected targets, obscuring true worst-case effectiveness. We argue that the right evaluation focuses on the hardest samples to poison. The same reasoning applies to defense: since targeted attacks leave no footprint at the distribution level, defenders should proactively identify the most vulnerable samples and apply targeted countermeasures. Given a test dataset, this paper identifies both the easiest and hardest to poison examples based on only clean model information. Specifically, we offer coarse evaluations using clean training dynamics, and fine-grained classification on poison class using poison distances and budgets. Our experiments show these metrics reliably stratify samples by poisoning vulnerability, enabling both rigorous worst-case evaluation and proactive vulnerability-aware defense.
翻译:目标投毒攻击通过向训练数据中注入恶意样本,操纵模型对特定测试样本的预测。然而,现有评估通常报告随机选取目标上的平均攻击成功率,这掩盖了真实的最坏情况效果。我们认为,正确的评估应聚焦于最难投毒的样本。同理适用于防御:由于目标攻击在分布层面未留下痕迹,防御者应主动识别最脆弱的样本,并采取针对性反制措施。针对给定测试数据集,本文仅利用干净模型信息,识别出最易与最难投毒的样本。具体而言,我们利用干净训练动态提供粗略评估,并基于投毒距离与预算对投毒类别进行细粒度分类。实验表明,这些指标能可靠地根据投毒脆弱性对样本进行分层,从而支持严谨的最坏情况评估与主动的脆弱性感知防御。