Having reliable specifications is an unavoidable challenge in achieving verifiable correctness, robustness, and interpretability of AI systems. Existing specifications for neural networks are in the paradigm of data as specification. That is, the local neighborhood centering around a reference input is considered to be correct (or robust). While existing specifications contribute to verifying adversarial robustness, a significant problem in many research domains, our empirical study shows that those verified regions are somewhat tight, and thus fail to allow verification of test set inputs, making them impractical for some real-world applications. To this end, we propose a new family of specifications called neural representation as specification, which uses the intrinsic information of neural networks - neural activation patterns (NAPs), rather than input data to specify the correctness and/or robustness of neural network predictions. We present a simple statistical approach to mining neural activation patterns. To show the effectiveness of discovered NAPs, we formally verify several important properties, such as various types of misclassifications will never happen for a given NAP, and there is no ambiguity between different NAPs. We show that by using NAP, we can verify a significant region of the input space, while still recalling 84% of the data on MNIST. Moreover, we can push the verifiable bound to 10 times larger on the CIFAR10 benchmark. Thus, we argue that NAPs can potentially be used as a more reliable and extensible specification for neural network verification.
翻译:拥有可靠的规范是实现AI系统可验证的正确性、鲁棒性和可解释性中不可避免的挑战。现有的神经网络规范遵循“数据即规范”的范式,即围绕参考输入的局部邻域被视为正确(或鲁棒)。尽管现有规范有助于验证对抗鲁棒性(这在许多研究领域是一个重要问题),但我们的实证研究表明,这些验证区域相对狭窄,导致无法验证测试集输入,使其在某些实际应用中不切实际。为此,我们提出了一类新的规范,称为“神经表征即规范”,它利用神经网络的固有信息——神经激活模式(NAPs),而非输入数据来指定神经网络预测的正确性和/或鲁棒性。我们提出了一种简单的统计方法来挖掘神经激活模式。为展示发现的NAPs的有效性,我们正式验证了几项重要属性,例如给定NAP下不会发生多种类型的误分类,且不同NAPs之间不存在歧义。我们证明,通过使用NAP,可以验证输入空间的显著区域,同时在MNIST上仍能召回84%的数据。此外,在CIFAR10基准测试上,我们可将可验证边界推大至10倍。因此,我们认为NAPs有潜力成为神经网络验证中更可靠、更可扩展的规范。