Adversarial training is a widely used method to improve the robustness of deep neural networks (DNNs) over adversarial perturbations. However, it is empirically observed that adversarial training on over-parameterized networks often suffers from the \textit{robust overfitting}: it can achieve almost zero adversarial training error while the robust generalization performance is not promising. In this paper, we provide a theoretical understanding of the question of whether overfitted DNNs in adversarial training can generalize from an approximation viewpoint. Specifically, our main results are summarized into three folds: i) For classification, we prove by construction the existence of infinitely many adversarial training classifiers on over-parameterized DNNs that obtain arbitrarily small adversarial training error (overfitting), whereas achieving good robust generalization error under certain conditions concerning the data quality, well separated, and perturbation level. ii) Linear over-parameterization (meaning that the number of parameters is only slightly larger than the sample size) is enough to ensure such existence if the target function is smooth enough. iii) For regression, our results demonstrate that there also exist infinitely many overfitted DNNs with linear over-parameterization in adversarial training that can achieve almost optimal rates of convergence for the standard generalization error. Overall, our analysis points out that robust overfitting can be avoided but the required model capacity will depend on the smoothness of the target function, while a robust generalization gap is inevitable. We hope our analysis will give a better understanding of the mathematical foundations of robustness in DNNs from an approximation view.
翻译:对抗训练是一种广泛用于提升深度神经网络(DNNs)对抗扰动鲁棒性的方法。然而,实证观察表明,过参数化网络上的对抗训练常遭受\textit{鲁棒过拟合}:它可实现几乎为零的对抗训练误差,但鲁棒泛化性能却不理想。本文从逼近视角对对抗训练中过拟合DNNs能否泛化的问题提供理论理解。具体而言,我们的主要结果归纳为三点:i)对于分类问题,我们通过构造证明,过参数化DNNs上存在无穷多个对抗训练分类器,可获得任意小的对抗训练误差(过拟合),同时在数据质量、良好分离性和扰动水平满足特定条件下实现良好的鲁棒泛化误差。ii)若目标函数足够光滑,线性过参数化(即参数数量仅略大于样本量)即可保证此类存在性。iii)对于回归问题,我们的结果表明,对抗训练中同样存在无穷多个线性过参数化的过拟合DNNs,它们能实现标准泛化误差的几乎最优收敛速率。总体而言,我们的分析指出鲁棒过拟合可以避免,但所需模型容量将取决于目标函数的光滑性,而鲁棒泛化差距不可避免。我们希望该分析能从逼近视角促进对DNNs鲁棒性数学基础的理解。