Shapley value attribution (SVA) is an increasingly popular explainable AI (XAI) method, which quantifies the contribution of each feature to the model's output. However, recent work has shown that most existing methods to implement SVAs have some drawbacks, resulting in biased or unreliable explanations that fail to correctly capture the true intrinsic relationships between features and model outputs. Moreover, the mechanism and consequences of these drawbacks have not been discussed systematically. In this paper, we propose a novel error theoretical analysis framework, in which the explanation errors of SVAs are decomposed into two components: observation bias and structural bias. We further clarify the underlying causes of these two biases and demonstrate that there is a trade-off between them. Based on this error analysis framework, we develop two novel concepts: over-informative and underinformative explanations. We demonstrate how these concepts can be effectively used to understand potential errors of existing SVA methods. In particular, for the widely deployed assumption-based SVAs, we find that they can easily be under-informative due to the distribution drift caused by distributional assumptions. We propose a measurement tool to quantify such a distribution drift. Finally, our experiments illustrate how different existing SVA methods can be over- or under-informative. Our work sheds light on how errors incur in the estimation of SVAs and encourages new less error-prone methods.
翻译:Shapley值归因(SVA)是一种日益流行的可解释人工智能(XAI)方法,它量化了每个特征对模型输出的贡献。然而,近期研究表明,大多数现有的SVA实现方法存在一些缺陷,导致产生有偏或不可靠的解释,未能正确捕捉特征与模型输出之间真实的内在关系。此外,这些缺陷的机制与后果尚未得到系统性的讨论。本文提出了一种新颖的误差理论分析框架,将SVA的解释误差分解为两个组成部分:观测偏差与结构偏差。我们进一步阐明了这两种偏差的根本成因,并证明二者之间存在权衡关系。基于此误差分析框架,我们提出了两个新概念:过度信息解释与信息不足解释。我们展示了如何有效利用这些概念来理解现有SVA方法的潜在误差。特别地,对于广泛部署的基于假设的SVA方法,我们发现由于分布假设引起的分布漂移,它们很容易产生信息不足的解释。我们提出了一种量化此类分布漂移的测量工具。最后,实验展示了不同现有SVA方法如何可能产生过度信息或信息不足的解释。本研究揭示了SVA估计中误差的产生机制,并为开发误差更少的新方法提供了启示。