In recent years, there has been increasing interest in explanation methods for neural model predictions that offer precise formal guarantees. These include abductive (respectively, contrastive) methods, which aim to compute minimal subsets of input features that are sufficient for a given prediction to hold (respectively, to change a given prediction). The corresponding decision problems are, however, known to be intractable. In this paper, we investigate whether tractability can be regained by focusing on neural models implementing a monotonic function. Although the relevant decision problems remain intractable, we can show that they become solvable in polynomial time by means of greedy algorithms if we additionally assume that the activation functions are continuous everywhere and differentiable almost everywhere. Our experiments suggest favourable performance of our algorithms.
翻译:近年来,对具备精确形式化保证的神经网络模型预测解释方法的研究兴趣日益增长。这些方法包括溯因解释(分别对应对比解释),旨在计算足以维持(分别改变)给定预测结果的最小输入特征子集。然而,相应的决策问题已知是难以处理的。本文探究通过聚焦实现单调函数的神经网络模型能否恢复可处理性。尽管相关决策问题仍然难以处理,但若附加假设激活函数处处连续且几乎处处可微,我们可证明此类问题可通过贪心算法在多项式时间内求解。实验表明所提算法具有良好的性能表现。