Analysis of how semantic concepts are represented within Convolutional Neural Networks (CNNs) is a widely used approach in Explainable Artificial Intelligence (XAI) for interpreting CNNs. A motivation is the need for transparency in safety-critical AI-based systems, as mandated in various domains like automated driving. However, to use the concept representations for safety-relevant purposes, like inspection or error retrieval, these must be of high quality and, in particular, stable. This paper focuses on two stability goals when working with concept representations in computer vision CNNs: stability of concept retrieval and of concept attribution. The guiding use-case is a post-hoc explainability framework for object detection (OD) CNNs, towards which existing concept analysis (CA) methods are successfully adapted. To address concept retrieval stability, we propose a novel metric that considers both concept separation and consistency, and is agnostic to layer and concept representation dimensionality. We then investigate impacts of concept abstraction level, number of concept training samples, CNN size, and concept representation dimensionality on stability. For concept attribution stability we explore the effect of gradient instability on gradient-based explainability methods. The results on various CNNs for classification and object detection yield the main findings that (1) the stability of concept retrieval can be enhanced through dimensionality reduction via data aggregation, and (2) in shallow layers where gradient instability is more pronounced, gradient smoothing techniques are advised. Finally, our approach provides valuable insights into selecting the appropriate layer and concept representation dimensionality, paving the way towards CA in safety-critical XAI applications.
翻译:卷积神经网络(CNN)中语义概念表示的分析是可解释人工智能(XAI)中解读CNN的广泛采用方法。其动机源于安全关键型AI系统对透明度的需求,这在自动驾驶等领域已被强制要求。然而,若要将概念表示用于检查或错误检索等安全相关目的,这些表示必须具有高质量,尤其是稳定性。本文聚焦于计算机视觉CNN中概念表示的两项稳定性目标:概念检索稳定性与概念归因稳定性。指导性用例是为目标检测(OD)CNN设计的后置可解释性框架,现有概念分析(CA)方法被成功适配至该场景。针对概念检索稳定性,我们提出一种新型度量指标,该指标同时考虑概念分离性与一致性,且不受层与概念表示维度的限制。随后探讨了概念抽象层级、概念训练样本数量、CNN模型规模及概念表示维度对稳定性的影响。针对概念归因稳定性,我们研究了梯度不稳定性对基于梯度的可解释性方法的影响。在多种用于分类与目标检测的CNN上的实验获得主要发现:(1)通过数据聚合进行降维可增强概念检索稳定性;(2)在梯度不稳定性更显著的浅层网络中,建议采用梯度平滑技术。最终,我们的方法为选择适当的层与概念表示维度提供了重要见解,为安全关键型XAI应用中的概念分析铺平道路。