Attribution methods are primarily designed to study the distribution of input component contributions to individual model predictions. However, some research applications require a summary of attribution patterns across the entire dataset to facilitate the interpretability of the scrutinized models. In this paper, we present a new method called Integrated Gradient Correlation (IGC) that relates dataset-wise attributions to a model prediction score and enables region-specific analysis by a direct summation over associated components. We demonstrate our method on scalar predictions with the study of image feature representation in the brain from fMRI neural signals and the estimation of neural population receptive fields (NSD dataset), as well as on categorical predictions with the investigation of handwritten digit recognition (MNIST dataset). The resulting IGC attributions show selective patterns, revealing underlying model strategies coherent with their respective objectives.
翻译:归因方法主要用于研究输入组件对单个模型预测的贡献分布。然而,一些研究应用需要总结整个数据集上的归因模式,以增强被检查模型的可解释性。本文提出了一种新方法——集成梯度相关性(IGC),该方法将数据集层面的归因与模型预测分数联系起来,并通过直接对相关组件求和实现区域特异性分析。我们在标量预测任务上展示了该方法,包括通过fMRI神经信号研究大脑中的图像特征表征以及神经群体感受野估计(NSD数据集),同时在分类预测任务上通过手写数字识别研究(MNIST数据集)进行了验证。得到的IGC归因展现出选择性模式,揭示了与各自目标相一致的底层模型策略。