Concept Bottleneck Models (CBMs) first map raw input(s) to a vector of human-defined concepts, before using this vector to predict a final classification. We might therefore expect CBMs capable of predicting concepts based on distinct regions of an input. In doing so, this would support human interpretation when generating explanations of the model's outputs to visualise input features corresponding to concepts. The contribution of this paper is threefold: Firstly, we expand on existing literature by looking at relevance both from the input to the concept vector, confirming that relevance is distributed among the input features, and from the concept vector to the final classification where, for the most part, the final classification is made using concepts predicted as present. Secondly, we report a quantitative evaluation to measure the distance between the maximum input feature relevance and the ground truth location; we perform this with the techniques, Layer-wise Relevance Propagation (LRP), Integrated Gradients (IG) and a baseline gradient approach, finding LRP has a lower average distance than IG. Thirdly, we propose using the proportion of relevance as a measurement for explaining concept importance.
翻译:概念瓶颈模型(CBM)首先将原始输入映射到人类定义的概念向量,再基于该向量预测最终分类结果。因此,我们预期CBM能够根据输入中的不同区域预测概念,从而在生成模型输出解释时,可通过可视化输入特征与对应概念的关系辅助人类理解。本文贡献有三:其一,我们扩展了现有文献的研究范畴,既从输入到概念向量的相关性分析(确认相关特征分布于输入各区域),又从概念向量到最终分类的相关性分析(发现多数情况下,最终分类主要依赖被预测为存在的概念);其二,我们提出量化评估方法,通过层间相关传播(LRP)、积分梯度(IG)和基线梯度方法测量最大输入特征相关性与真实标注位置的距离,实验表明LRP的平均距离低于IG;其三,我们提出将相关性比例作为衡量概念重要性的量化指标。