The robustness of deep neural networks (DNNs) is crucial to the hosting system's reliability and security. Formal verification has been demonstrated to be effective in providing provable robustness guarantees. To improve its scalability, over-approximating the non-linear activation functions in DNNs by linear constraints has been widely adopted, which transforms the verification problem into an efficiently solvable linear programming problem. Many efforts have been dedicated to defining the so-called tightest approximations to reduce overestimation imposed by over-approximation. In this paper, we study existing approaches and identify a dominant factor in defining tight approximation, namely the approximation domain of the activation function. We find out that tight approximations defined on approximation domains may not be as tight as the ones on their actual domains, yet existing approaches all rely only on approximation domains. Based on this observation, we propose a novel dual-approximation approach to tighten over-approximations, leveraging an activation function's underestimated domain to define tight approximation bounds. We implement our approach with two complementary algorithms based respectively on Monte Carlo simulation and gradient descent into a tool called DualApp. We assess it on a comprehensive benchmark of DNNs with different architectures. Our experimental results show that DualApp significantly outperforms the state-of-the-art approaches with 100% - 1000% improvement on the verified robustness ratio and 10.64% on average (up to 66.53%) on the certified lower bound.
翻译:深度神经网络(DNN)的鲁棒性对其所在系统的可靠性与安全性至关重要。形式化验证已被证明能够有效提供可证明的鲁棒性保证。为提升其可扩展性,通过线性约束对DNN中的非线性激活函数进行过逼近已成为广泛采用的方法,该方法可将验证问题转化为高效可解的线性规划问题。已有大量工作致力于定义所谓的最紧逼近,以减少过逼近引入的过度估计。本文研究现有方法,并识别出定义紧逼近的一个主导因素——激活函数的逼近域。我们发现,在逼近域上定义的紧逼近可能不如在实际域上定义的紧逼近,然而现有方法均仅依赖逼近域。基于这一观察,我们提出一种新颖的双重逼近方法以收紧过逼近,通过利用激活函数的欠估计域来定义紧逼近边界。我们分别基于蒙特卡洛模拟和梯度下降实现两种互补算法,并将其集成到名为DualApp的工具中。我们在涵盖不同架构DNN的综合基准上进行了评估。实验结果表明,DualApp显著优于现有最优方法,在验证鲁棒性比率上提升100%至1000%,在认证下界上平均提升10.64%(最高可达66.53%)。