The streams of research on adversarial examples and counterfactual explanations have largely been growing independently. This has led to several recent works trying to elucidate their similarities and differences. Most prominently, it has been argued that adversarial examples, as opposed to counterfactual explanations, have a unique characteristic in that they lead to a misclassification compared to the ground truth. However, the computational goals and methodologies employed in existing counterfactual explanation and adversarial example generation methods often lack alignment with this requirement. Using formal definitions of adversarial examples and counterfactual explanations, we introduce non-adversarial algorithmic recourse and outline why in high-stakes situations, it is imperative to obtain counterfactual explanations that do not exhibit adversarial characteristics. We subsequently investigate how different components in the objective functions, e.g., the machine learning model or cost function used to measure distance, determine whether the outcome can be considered an adversarial example or not. Our experiments on common datasets highlight that these design choices are often more critical in deciding whether recourse is non-adversarial than whether recourse or attack algorithms are used. Furthermore, we show that choosing a robust and accurate machine learning model results in less adversarial recourse desired in practice.
翻译:对抗样本与反事实解释的研究领域在很大程度上是独立发展的。这导致近期多项工作尝试阐明它们之间的异同。最显著的是,有观点认为对抗样本与反事实解释相比具有独特特征,即它们会导致与真实结果相比的错误分类。然而,现有反事实解释和对抗样本生成方法在计算目标和方法论上常与这一要求不符。通过使用对抗样本和反事实解释的形式化定义,我们引入了非对抗性算法救济,并概述了为何在高风险情境下,获得不具有对抗性特征的反事实解释至关重要。随后,我们探讨了目标函数中的不同组成部分(例如,用于测量距离的机器学习模型或代价函数)如何决定结果是否可被视为对抗样本。我们在常见数据集上的实验表明,相较于使用救济算法还是攻击算法,这些设计选择往往更关键地决定了救济是否具有非对抗性。此外,我们证明选择鲁棒且准确的机器学习模型能够产生实践中更符合需求、对抗性更弱的救济。