Post-hoc explanation methods are an important tool for increasing model transparency for users. Unfortunately, the currently used methods for attributing token importance often yield diverging patterns. In this work, we study potential sources of disagreement across methods from a linguistic perspective. We find that different methods systematically select different classes of words and that methods that agree most with other methods and with humans display similar linguistic preferences. Token-level differences between methods are smoothed out if we compare them on the syntactic span level. We also find higher agreement across methods by estimating the most important spans dynamically instead of relying on a fixed subset of size $k$. We systematically investigate the interaction between $k$ and spans and propose an improved configuration for selecting important tokens.
翻译:事后解释方法是增强模型对用户透明度的重要工具。然而,当前用于归因词元重要性的方法往往会产生分歧模式。本研究从语言学视角探讨不同方法间潜在的分歧来源。我们发现,不同方法会系统地选择不同类别的词语,且与其他方法及人类判断一致性较高的方法呈现出相似的语言偏好。若在句法跨度层面对比不同方法,词元层面的差异会被平滑处理。我们还发现,通过动态估计最重要的跨度而非依赖固定大小为$k$的子集,方法间的一致性会更高。我们系统研究了$k$与跨度之间的交互作用,并提出了一种用于选择重要词元的改进配置方案。