Layer-wise relevance propagation (LRP) is a widely used and powerful technique to reveal insights into various artificial neural network (ANN) architectures. LRP is often used in the context of image classification. The aim is to understand, which parts of the input sample have highest relevance and hence most influence on the model prediction. Relevance can be traced back through the network to attribute a certain score to each input pixel. Relevance scores are then combined and displayed as heat maps and give humans an intuitive visual understanding of classification models. Opening the black box to understand the classification engine in great detail is essential for domain experts to gain trust in ANN models. However, there are pitfalls in terms of model-inherent artifacts included in the obtained relevance maps, that can easily be missed. But for a valid interpretation, these artifacts must not be ignored. Here, we apply and revise LRP on various ANN architectures trained as classifiers on geospatial and synthetic data. Depending on the network architecture, we show techniques to control model focus and give guidance to improve the quality of obtained relevance maps to separate facts from artifacts.
翻译:逐层相关性传播(LRP)是一种广泛应用于揭示各类人工神经网络(ANN)架构内部机制的强大技术。LRP常用于图像分类场景,其目标是理解输入样本的哪些部分具有最高相关性,从而对模型预测影响最大。通过沿网络反向追溯相关性,可为每个输入像素分配特定分数。这些分数经组合后以热力图形式呈现,为人类提供对分类模型的直观视觉理解。打开黑箱以深入理解分类引擎,对于领域专家建立对ANN模型的信任至关重要。然而,在获取的相关性图中存在模型固有伪影的陷阱,这些伪影极易被忽略,但为获得有效解释,绝不能忽视它们。本文针对基于地理空间与合成数据训练的分类器,在多种ANN架构上应用并修订了LRP方法。我们根据不同网络架构展示了控制模型聚焦的技术,并提出了提升相关性图质量的指导原则,以区分事实与伪影。