Fast Adversarial Training (FAT) has attracted significant attention due to its efficiency in enhancing neural network robustness against adversarial attacks. However, FAT is prone to catastrophic overfitting (CO), wherein models overfit to the specific attack used during training and fail to generalize to others. While existing methods introduce diverse hypotheses and propose various strategies to mitigate CO, a systematic and intuitive explanation of CO remains absent. In this work, we innovatively interpret CO through the lens of backdoor. Through validations on pathway division, diverse feature predictions, and universal class distinguishable triggers in CO, we conceptualize CO as a weak trigger variant of unlearnable tasks, unifying CO, backdoor attacks, and unlearnable tasks under a common theoretical framework. Guided by this, we leverage several backdoor inspired strategies to mitigate CO: (i) Recalibrate CO affected model parameters using vanilla fine tuning, linear probing, or reinitialization-based techniques; (ii) Introduce a weight outlier suppression constraint to regulate abnormal deviations in model weights. Extensive experiments support our interpretation of CO and show the efficacy of the proposed mitigation strategies.
翻译:快速对抗训练(FAT)因其在提升神经网络对抗攻击鲁棒性方面的高效性而备受关注。然而,FAT易出现灾难性过拟合(CO),即模型过度拟合训练中使用的特定攻击,无法泛化至其他攻击类型。现有方法虽提出多种假设并制定缓解CO的策略,但仍缺乏对CO的系统性直观解释。本文创新性地从后门视角解读CO。通过验证CO中的路径划分、多样特征预测及通用类别可区分触发器,我们将CO视为不可学习任务的弱触发变体,从而将CO、后门攻击与不可学习任务统一至共同理论框架。基于此,我们采用多种后门启发式策略缓解CO:(i) 通过微调、线性探测或重初始化技术重新校准受CO影响的模型参数;(ii) 引入权重异常值抑制约束以调控模型权重的异常偏差。大量实验支持我们对CO的解释,并证明了所提缓解策略的有效性。