Counterfactual explanations play an important role in detecting bias and improving the explainability of data-driven classification models. A counterfactual explanation (CE) is a minimal perturbed data point for which the decision of the model changes. Most of the existing methods can only provide one CE, which may not be achievable for the user. In this work we derive an iterative method to calculate robust CEs, i.e. CEs that remain valid even after the features are slightly perturbed. To this end, our method provides a whole region of CEs allowing the user to choose a suitable recourse to obtain a desired outcome. We use algorithmic ideas from robust optimization and prove convergence results for the most common machine learning methods including logistic regression, decision trees, random forests, and neural networks. Our experiments show that our method can efficiently generate globally optimal robust CEs for a variety of common data sets and classification models.
翻译:反事实解释在检测偏差和提升数据驱动分类模型的可解释性方面发挥着重要作用。反事实解释(CE)是指经过最小扰动后模型决策发生改变的数据点。现有方法大多只能提供单一反事实解释,这在实际应用中可能难以被用户实现。我们提出了一种迭代方法,用于计算鲁棒反事实解释,即即使特征发生微小扰动仍保持有效性的解释。为此,我们的方法能够生成包含多个反事实解释的区域,使用户可以从中选择合适的行动路径以获得期望结果。我们借鉴鲁棒优化的算法思想,针对逻辑回归、决策树、随机森林和神经网络等最常见机器学习方法证明了收敛性结果。实验表明,该方法能高效地为多种常见数据集和分类模型生成全局最优的鲁棒反事实解释。