While enjoying the great achievements brought by deep learning (DL), people are also worried about the decision made by DL models, since the high degree of non-linearity of DL models makes the decision extremely difficult to understand. Consequently, attacks such as adversarial attacks are easy to carry out, but difficult to detect and explain, which has led to a boom in the research on local explanation methods for explaining model decisions. In this paper, we evaluate the faithfulness of explanation methods and find that traditional tests on faithfulness encounter the random dominance problem, \ie, the random selection performs the best, especially for complex data. To further solve this problem, we propose three trend-based faithfulness tests and empirically demonstrate that the new trend tests can better assess faithfulness than traditional tests on image, natural language and security tasks. We implement the assessment system and evaluate ten popular explanation methods. Benefiting from the trend tests, we successfully assess the explanation methods on complex data for the first time, bringing unprecedented discoveries and inspiring future research. Downstream tasks also greatly benefit from the tests. For example, model debugging equipped with faithful explanation methods performs much better for detecting and correcting accuracy and security problems.
翻译:尽管深度学习取得了巨大成就,但人们也对其模型做出的决策感到担忧,因为深度学习模型的高度非线性使得决策极难理解。因此,对抗攻击等攻击易于实施,却难以检测和解释,这导致针对解释模型决策的局部解释方法研究蓬勃发展。本文评估了解释方法的忠实性,并发现传统忠实性测试存在随机主导问题,即随机选择表现最佳,尤其是在复杂数据上。为解决该问题,我们提出了三种基于趋势的忠实性测试,并通过实验证明,在图像、自然语言和安全任务中,新趋势测试能比传统测试更准确地评估忠实性。我们实现了评估系统,并对十种流行解释方法进行了评估。借助趋势测试,我们首次成功评估了复杂数据上的解释方法,带来了前所未见的发现,为未来研究提供了启示。下游任务也从这些测试中获益良多。例如,配备忠实解释方法的模型调试在检测和纠正精度及安全问题方面表现更佳。