The internal functional behavior of trained Deep Neural Networks is notoriously difficult to interpret. Activation-maximization approaches are one set of techniques used to interpret and analyze trained deep-learning models. These consist in finding inputs that maximally activate a given neuron or feature map. These inputs can be selected from a data set or obtained by optimization. However, interpretability methods may be subject to being deceived. In this work, we consider the concept of an adversary manipulating a model for the purpose of deceiving the interpretation. We propose an optimization framework for performing this manipulation and demonstrate a number of ways that popular activation-maximization interpretation techniques associated with CNNs can be manipulated to change the interpretations, shedding light on the reliability of these methods.
翻译:训练有素的深度神经网络的内部功能性行为以难以解释著称。激活最大化方法是一类用于解释和分析训练后深度学习模型的技术,其核心在于寻找能最大程度激活特定神经元或特征图的输入。这些输入可从数据集中选取,也可通过优化获得。然而,可解释性方法可能面临被欺骗的风险。本研究探讨了对手操纵模型以欺骗解释方法的概念,提出了一种实现这种操纵的优化框架,并展示了多种操纵与卷积神经网络关联的流行激活最大化解释技术以改变其解释结果的方式,揭示了这些方法的可靠性问题。