The evaluation of explainable artificial intelligence is challenging, because automated and human-centred metrics of explanation quality may diverge. To clarify their relationship, we investigated whether human and artificial image classification will benefit from the same visual explanations. In three experiments, we analysed human reaction times, errors, and subjective ratings while participants classified image segments. These segments either reflected human attention (eye movements, manual selections) or the outputs of two attribution methods explaining a ResNet (Grad-CAM, XRAI). We also had this model classify the same segments. Humans and the model largely agreed on the interpretability of attribution methods: Grad-CAM was easily interpretable for indoor scenes and landscapes, but not for objects, while the reverse pattern was observed for XRAI. Conversely, human and model performance diverged for human-generated segments. Our results caution against general statements about interpretability, as it varies with the explanation method, the explained images, and the agent interpreting them.
翻译:可解释人工智能的评估颇具挑战性,因为解释质量的自动化指标与以人为中心的指标可能产生分歧。为厘清二者关系,我们通过三项实验探究了人类与人工智能图像分类是否受益于相同的视觉解释。实验中,参与者在分类图像片段时,我们记录了其反应时间、错误率及主观评分。这些图像片段或反映人类注意力(眼动轨迹、手动选择),或反映解释ResNet的两种归因方法(Grad-CAM、XRAI)的输出结果。我们还让同一模型对这些片段进行分类。人类与模型在归因方法的可解释性上基本达成一致:Grad-CAM对室内场景和风景图像易于解释,但对物体图像效果欠佳;而XRAI则呈现相反模式。相反,对于人类生成的片段,人类与模型的性能表现存在分歧。我们的研究结果警示,不应笼统定义可解释性,因其随解释方法、被解释图像及解释主体的不同而变化。