In recent years numerous methods have been developed to formally verify the robustness of deep neural networks (DNNs). Though the proposed techniques are effective in providing mathematical guarantees about the DNNs behavior, it is not clear whether the proofs generated by these methods are human-interpretable. In this paper, we bridge this gap by developing new concepts, algorithms, and representations to generate human understandable interpretations of the proofs. Leveraging the proposed method, we show that the robustness proofs of standard DNNs rely on spurious input features, while the proofs of DNNs trained to be provably robust filter out even the semantically meaningful features. The proofs for the DNNs combining adversarial and provably robust training are the most effective at selectively filtering out spurious features as well as relying on human-understandable input features.
翻译:近年来,众多方法被开发用于形式化验证深度神经网络(DNN)的鲁棒性。尽管这些技术能有效提供关于DNN行为的数学保证,但此类方法生成的证明是否具有人类可解释性尚不明确。本文通过开发新概念、算法和表示方法,弥合了这一差距,生成人类可理解的证明解读。利用所提出的方法,我们发现标准DNN的鲁棒性证明依赖于虚假输入特征,而经过可证明鲁棒性训练的DNN证明甚至过滤掉了语义上有意义的特征。结合对抗训练与可证明鲁棒性训练的DNN,其证明在选择性过滤虚假特征以及依赖人类可理解的输入特征方面最为有效。