Adversarial training is well-known to produce high-quality neural network models that are empirically robust against adversarial perturbations. Nevertheless, once a model has been adversarially trained, one often desires a certification that the model is truly robust against all future attacks. Unfortunately, when faced with adversarially trained models, all existing approaches have significant trouble making certifications that are strong enough to be practically useful. Linear programming (LP) techniques in particular face a "convex relaxation barrier" that prevent them from making high-quality certifications, even after refinement with mixed-integer linear programming (MILP) and branch-and-bound (BnB) techniques. In this paper, we propose a nonconvex certification technique, based on a low-rank restriction of a semidefinite programming (SDP) relaxation. The nonconvex relaxation makes strong certifications comparable to much more expensive SDP methods, while optimizing over dramatically fewer variables comparable to much weaker LP methods. Despite nonconvexity, we show how off-the-shelf local optimization algorithms can be used to achieve and to certify global optimality in polynomial time. Our experiments find that the nonconvex relaxation almost completely closes the gap towards exact certification of adversarially trained models.
翻译:对抗训练以生成在经验上对对抗扰动具有鲁棒性的高质量神经网络模型而闻名。然而,一旦模型经过对抗训练,人们通常希望获得该模型对将来所有攻击真正鲁棒的认证。不幸的是,面对对抗训练模型时,所有现有方法在生成足够强大以实际应用的认证时都遇到重大困难。特别是线性规划(LP)技术面临一个“凸松弛障碍”,即使通过混合整数线性规划(MILP)和分支定界(BnB)技术进行改进,也无法做出高质量的认证。在本文中,我们提出了一种基于半定规划(SDP)松弛的低秩限制的非凸认证技术。非凸松弛使得能够做出与更昂贵的SDP方法相媲美的强有力认证,同时优化的变量显著减少,与效果弱得多的LP方法相当。尽管非凸性,我们展示了如何利用现成的局部优化算法在多项式时间内实现并认证全局最优性。我们的实验发现,非凸松弛几乎完全弥合了对抗训练模型精确认证的差距。