Adversarial training is a standard method to train deep neural networks to be robust to adversarial perturbation. Similar to surprising $\textit{clean generalization}$ ability in the standard deep learning setting, neural networks trained by adversarial training also generalize well for $\textit{unseen clean data}$. However, in constrast with clean generalization, while adversarial training method is able to achieve low $\textit{robust training error}$, there still exists a significant $\textit{robust generalization gap}$, which promotes us exploring what mechanism leads to both $\textit{clean generalization and robust overfitting (CGRO)}$ during learning process. In this paper, we provide a theoretical understanding of this CGRO phenomenon in adversarial training. First, we propose a theoretical framework of adversarial training, where we analyze $\textit{feature learning process}$ to explain how adversarial training leads network learner to CGRO regime. Specifically, we prove that, under our patch-structured dataset, the CNN model provably partially learns the true feature but exactly memorizes the spurious features from training-adversarial examples, which thus results in clean generalization and robust overfitting. For more general data assumption, we then show the efficiency of CGRO classifier from the perspective of $\textit{representation complexity}$. On the empirical side, to verify our theoretical analysis in real-world vision dataset, we investigate the $\textit{dynamics of loss landscape}$ during training. Moreover, inspired by our experiments, we prove a robust generalization bound based on $\textit{global flatness}$ of loss landscape, which may be an independent interest.
翻译:对抗训练是训练深度神经网络以抵御对抗性扰动的标准方法。与标准深度学习设置中令人惊讶的$\textit{干净泛化}$能力类似,通过对抗训练训练的神经网络也能很好地泛化到$\textit{未见过的干净数据}$。然而,与干净泛化形成对比的是,尽管对抗训练方法能够实现较低的$\textit{鲁棒训练误差}$,但仍存在显著的$\textit{鲁棒泛化差距}$,这促使我们探索在学习过程中导致$\textit{干净泛化与鲁棒过拟合(CGRO)}$同时发生的机制。本文从理论上理解了对抗训练中的CGRO现象。首先,我们提出了一个对抗训练的理论框架,通过分析$\textit{特征学习过程}$来阐释对抗训练如何引导网络学习器进入CGRO状态。具体地,我们证明,在我们提出的块状结构数据集下,CNN模型被证明会部分学习真实特征但精确记忆训练-对抗样本中的伪特征,从而产生干净泛化和鲁棒过拟合。针对更一般的数据假设,我们从$\textit{表示复杂度}$角度展示了CGRO分类器的效率。在实证方面,为了验证我们在真实世界视觉数据集上的理论分析,我们研究了训练过程中$\textit{损失景观的动态变化}$。此外,受实验启发,我们证明了基于损失景观$\textit{全局平坦度}$的鲁棒泛化界,这可能具有独立的研究价值。