Machine Learning (ML) algorithms are vulnerable to poisoning attacks, where a fraction of the training data is manipulated to deliberately degrade the algorithms' performance. Optimal attacks can be formulated as bilevel optimization problems and help to assess their robustness in worst-case scenarios. We show that current approaches, which typically assume that hyperparameters remain constant, lead to an overly pessimistic view of the algorithms' robustness and of the impact of regularization. We propose a novel optimal attack formulation that considers the effect of the attack on the hyperparameters and models the attack as a multiobjective bilevel optimization problem. This allows to formulate optimal attacks, learn hyperparameters and evaluate robustness under worst-case conditions. We apply this attack formulation to several ML classifiers using $L_2$ and $L_1$ regularization. Our evaluation on multiple datasets confirms the limitations of previous strategies and evidences the benefits of using $L_2$ and $L_1$ regularization to dampen the effect of poisoning attacks.
翻译:机器学习算法容易受到投毒攻击,即训练数据中的部分样本被恶意操纵以故意降低算法性能。最优攻击可被形式化为双层优化问题,并有助于评估算法在最坏情况下的鲁棒性。我们表明,当前通常假设超参数保持不变的攻击建模方法会导致对算法鲁棒性及正则化影响的过度悲观评估。为此,我们提出一种考虑攻击对超参数影响的新型最优攻击建模方法,将攻击形式化为多目标双层优化问题。该框架能够制定最优攻击策略、学习超参数并在最坏条件下评估鲁棒性。我们将此攻击建模方法应用于采用 $L_2$ 和 $L_1$ 正则化的多个机器学习分类器。在多个数据集上的评估结果证实了先前策略的局限性,并揭示了使用 $L_2$ 和 $L_1$ 正则化可有效抑制投毒攻击影响的双重优势。