Various attribution methods have been developed to explain deep neural networks (DNNs) by inferring the attribution/importance/contribution score of each input variable to the final output. However, existing attribution methods are often built upon different heuristics. There remains a lack of a unified theoretical understanding of why these methods are effective and how they are related. To this end, for the first time, we formulate core mechanisms of fourteen attribution methods, which were designed on different heuristics, into the same mathematical system, i.e., the system of Taylor interactions. Specifically, we prove that attribution scores estimated by fourteen attribution methods can all be reformulated as the weighted sum of two types of effects, i.e., independent effects of each individual input variable and interaction effects between input variables. The essential difference among the fourteen attribution methods mainly lies in the weights of allocating different effects. Based on the above findings, we propose three principles for a fair allocation of effects to evaluate the faithfulness of the fourteen attribution methods.
翻译:各种归因方法已被开发用于解释深度神经网络(DNNs),通过推断每个输入变量对最终输出的归因/重要性/贡献分数。然而,现有的归因方法通常基于不同的启发式原则构建,目前仍缺乏对这些方法为何有效以及它们之间如何关联的统一理论理解。为此,我们首次将十四种基于不同启发式原则设计的归因方法的核心机制纳入同一数学系统,即泰勒交互作用系统。具体来说,我们证明了这十四种归因方法估计的归因分数均可重构为两类效应的加权和:每个输入变量的独立效应以及输入变量之间的交互效应。这十四种归因方法之间的本质差异主要在于分配不同效应时的权重不同。基于上述发现,我们提出了公平分配效应的三项原则,以评估这十四种归因方法的忠实度。