Part-prototype networks (e.g., ProtoPNet, ProtoTree, and ProtoPool) have attracted broad research interest for their intrinsic interpretability and comparable accuracy to non-interpretable counterparts. However, recent works find that the interpretability from prototypes is fragile, due to the semantic gap between the similarities in the feature space and that in the input space. In this work, we strive to address this challenge by making the first attempt to quantitatively and objectively evaluate the interpretability of the part-prototype networks. Specifically, we propose two evaluation metrics, termed as consistency score and stability score, to evaluate the explanation consistency across images and the explanation robustness against perturbations, respectively, both of which are essential for explanations taken into practice. Furthermore, we propose an elaborated part-prototype network with a shallow-deep feature alignment (SDFA) module and a score aggregation (SA) module to improve the interpretability of prototypes. We conduct systematical evaluation experiments and provide substantial discussions to uncover the interpretability of existing part-prototype networks. Experiments on three benchmarks across nine architectures demonstrate that our model achieves significantly superior performance to the state of the art, in both the accuracy and interpretability. Our code is available at https://github.com/hqhQAQ/EvalProtoPNet.
翻译:部分原型网络(如ProtoPNet、ProtoTree和ProtoPool)因其内在的可解释性以及与不可解释模型相当的准确率而受到广泛研究关注。然而,最新研究发现,由于特征空间相似性与输入空间相似性之间存在语义鸿沟,原型网络的可解释性较为脆弱。本研究首次尝试定量且客观地评估部分原型网络的可解释性,以应对这一挑战。具体而言,我们提出两个评估指标——一致性分数和稳定性分数,分别用于评估跨图像的解释一致性和对抗扰动的解释鲁棒性,这两者在实际应用的解释场景中均至关重要。此外,我们提出了一种精细化的部分原型网络,通过引入浅层-深层特征对齐模块和分数聚合模块来提升原型网络的可解释性。我们进行了系统性评估实验,并通过充分讨论揭示了现有部分原型网络的可解释性特征。在九个架构、三个基准数据集上的实验表明,我们的模型在准确率和可解释性方面均显著优于现有最佳方法。相关代码已开源至 https://github.com/hqhQAQ/EvalProtoPNet。