Modern deep learning models are notoriously opaque, which has motivated the development of methods for interpreting how deep models predict. This goal is usually approached with attribution method, which assesses the influence of features on model predictions. As an explanation method, the evaluation criteria of attribution methods is how accurately it re-reflects the actual reasoning process of the model (faithfulness). Meanwhile, since the reasoning process of deep models is inaccessible, researchers design various evaluation methods to demonstrate their arguments. However, some crucial logic traps in these evaluation methods are ignored in most works, causing inaccurate evaluation and unfair comparison. This paper systematically reviews existing methods for evaluating attribution scores and summarizes the logic traps in these methods. We further conduct experiments to demonstrate the existence of each logic trap. Through both the theoretical and experimental analysis, we hope to increase attention on the inaccurate evaluation of attribution scores. Moreover, with this paper, we suggest stopping focusing on improving performance under unreliable evaluation systems and starting efforts on reducing the impact of proposed logic traps
翻译:现代深度学习模型因其不透明性而闻名,这推动了深度模型预测解释方法的发展。这一目标通常通过归因方法实现,该方法评估特征对模型预测的影响。作为解释方法,归因方法的评估标准在于其能否准确反映模型的实际推理过程(忠实性)。同时,由于深度模型的推理过程不可访问,研究人员设计了各种评估方法来论证其观点。然而,这些评估方法中存在的一些关键逻辑陷阱在大多数工作中被忽视,导致评估不准确和比较不公平。本文系统性地回顾了现有的归因分数评估方法,并总结了这些方法中的逻辑陷阱。我们进一步通过实验证明了每个逻辑陷阱的存在。通过理论和实验分析,我们希望引起对归因分数不准确评估的更多关注。此外,本文建议停止在不可靠评估体系下追求性能提升,转而致力于减少所提出的逻辑陷阱的影响。