As the use of deep neural networks continues to grow, understanding their behaviour has become more crucial than ever. Post-hoc explainability methods are a potential solution, but their reliability is being called into question. Our research investigates the response of post-hoc visual explanations to naturally occurring transformations, often referred to as augmentations. We anticipate explanations to be invariant under certain transformations, such as changes to the colour map while responding in an equivariant manner to transformations like translation, object scaling, and rotation. We have found remarkable differences in robustness depending on the type of transformation, with some explainability methods (such as LRP composites and Guided Backprop) being more stable than others. We also explore the role of training with data augmentation. We provide evidence that explanations are typically less robust to augmentation than classification performance, regardless of whether data augmentation is used in training or not.
翻译:随着深度神经网络的持续应用,理解其行为变得比以往更加关键。事后可解释性方法是一种潜在解决方案,但其可靠性正受到质疑。本研究探究了事后视觉解释对自然发生的变换(通常称为数据增强)的响应。我们预期解释在特定变换(如颜色映射变化)下应具有不变性,而对平移、物体缩放和旋转等变换则应以等变方式响应。我们发现不同变换类型在鲁棒性上存在显著差异,某些可解释性方法(如LRP复合方法和引导反向传播)比其他方法更为稳定。我们还探讨了使用数据增强进行训练的作用。研究表明,无论训练中是否使用数据增强,解释对增强的鲁棒性通常低于分类性能。