Deep neural networks are being increasingly implemented throughout society in recent years. It is useful to identify which parameters trigger misclassification in diagnosing undesirable model behaviors. The concept of parameter saliency is proposed and used to diagnose convolutional neural networks (CNNs) by ranking convolution filters that may have caused misclassification on the basis of parameter saliency. It is also shown that fine-tuning the top ranking salient filters has efficiently corrected misidentification on ImageNet. However, there is still a knowledge gap in terms of understanding why parameter saliency ranking can find the filters inducing misidentification. In this work, we attempt to bridge the gap by analyzing parameter saliency ranking from a statistical viewpoint, namely, extreme value theory. We first show that the existing work implicitly assumes that the gradient norm computed for each filter follows a normal distribution. Then, we clarify the relationship between parameter saliency and the score based on the peaks-over-threshold (POT) method, which is often used to model extreme values. Finally, we reformulate parameter saliency in terms of the POT method, where this reformulation is regarded as statistical anomaly detection and does not require the implicit assumptions of the existing parameter-saliency formulation. Our experimental results demonstrate that our reformulation can detect malicious filters as well. Furthermore, we show that the existing parameter saliency method exhibits a bias against the depth of layers in deep neural networks. In particular, this bias has the potential to inhibit the discovery of filters that cause misidentification in situations where domain shift occurs. In contrast, parameter saliency based on POT shows less of this bias.
翻译:近年来,深度神经网络在社会各领域的应用日益广泛。为诊断不良模型行为,识别引发误分类的特定参数具有重要意义。本文提出参数显著性概念,通过基于参数显著性对可能引发误分类的卷积滤波器进行排序,用于诊断卷积神经网络。研究表明,对排名靠前的显著滤波器进行微调,能够有效纠正ImageNet上的误识别问题。然而,关于参数显著性排序为何能定位引发误分类的滤波器,仍存在认知空白。本文尝试从统计视角——即极值理论——分析参数显著性排序来填补这一空白。首先,我们证明现有方法隐含假设每个滤波器计算的梯度范数服从正态分布;其次,阐明参数显著性与常用于极值建模的峰值超阈值方法所得分数之间的关系;最终,基于峰值超阈值方法重构参数显著性公式,该重构方法被视为统计异常检测,无需现有参数显著性公式的隐含假设。实验结果表明,我们的重构方法同样能有效检测恶意滤波器。此外,我们指出现有参数显著性方法对深度神经网络层次深度存在偏差,这种偏差在领域偏移发生时可能阻碍发现导致误分类的滤波器。相比之下,基于峰值超阈值的参数显著性方法表现出更小的此类偏差。