Graph neural networks (GNNs) are highly effective on a variety of graph-related tasks; however, they lack interpretability and transparency. Current explainability approaches are typically local and treat GNNs as black-boxes. They do not look inside the model, inhibiting human trust in the model and explanations. Motivated by the ability of neurons to detect high-level semantic concepts in vision models, we perform a novel analysis on the behaviour of individual GNN neurons to answer questions about GNN interpretability, and propose new metrics for evaluating the interpretability of GNN neurons. We propose a novel approach for producing global explanations for GNNs using neuron-level concepts to enable practitioners to have a high-level view of the model. Specifically, (i) to the best of our knowledge, this is the first work which shows that GNN neurons act as concept detectors and have strong alignment with concepts formulated as logical compositions of node degree and neighbourhood properties; (ii) we quantitatively assess the importance of detected concepts, and identify a trade-off between training duration and neuron-level interpretability; (iii) we demonstrate that our global explainability approach has advantages over the current state-of-the-art -- we can disentangle the explanation into individual interpretable concepts backed by logical descriptions, which reduces potential for bias and improves user-friendliness.
翻译:图神经网络(GNN)在多种图相关任务中表现出色,但其可解释性与透明度不足。现有解释方法通常局限于局部视角,将GNN视为黑箱模型,未能深入模型内部结构,这削弱了人类对模型及其解释结果的信任。受视觉模型中神经元可检测高层语义概念的启发,我们对单个GNN神经元的行为进行了创新性分析,以回答GNN可解释性相关问题,并提出了评估GNN神经元可解释性的新指标。我们提出了一种基于神经元层面概念的全局解释方法,使实践者能够获得模型的高层认知。具体而言:(i)据我们所知,这是首个证明GNN神经元充当概念检测器,并与节点度、邻域特性等逻辑组合而成的概念存在强对齐的研究工作;(ii)我们定量评估了检测概念的重要性,并发现训练时长与神经元可解释性之间存在权衡;(iii)我们证明,与当前最先进方法相比,本全局解释方法具有显著优势:可将解释结果拆解为多个独立可解释概念,并为每个概念提供逻辑描述支持,从而降低潜在偏差并提升用户友好性。