Fast Adversarial Training (FAT) has proven effective in enhancing model robustness by encouraging networks to learn perturbation-invariant representations. However, FAT often suffers from catastrophic overfitting (CO), where the model overfits to the training attack and fails to generalize to unseen ones. Moreover, robustness oriented optimization typically leads to notable performance degradation on clean inputs, and such degradation becomes increasingly severe as the perturbation budget grows. In this work, we conduct a comprehensive analysis of how guidance strength affects model performance by modulating perturbation and supervision levels across distinct confidence groups. The findings reveal that low confidence samples are the primary contributors to CO and the robustness accuracy trade off. Building on this insight, we propose a Distribution-aware Dynamic Guidance (DDG) strategy that dynamically adjusts both the perturbation budget and supervision signal. Specifically, DDG scales the perturbation magnitude according to the sample confidence at the ground truth class, thereby guiding samples toward consistent decision boundaries while mitigating the influence of learning spurious correlations. Simultaneously, it dynamically adjusts the supervision signal based on the prediction state of each sample, preventing overemphasis on incorrect signals. To alleviate potential gradient instability arising from dynamic guidance, we further design a weighted regularization constraint. Extensive experiments on standard benchmarks demonstrate that DDG effectively alleviates both CO and the robustness accuracy trade off.
翻译:快速对抗训练(FAT)通过鼓励网络学习对抗扰动不变的表示,已被证明能有效增强模型鲁棒性。然而,FAT常遭受灾难性过拟合(CO)问题,即模型过度拟合训练攻击而无法泛化至未见攻击。此外,面向鲁棒性的优化通常会导致干净输入上的显著性能退化,且这种退化随扰动预算增加而日益严重。本文通过调节不同置信度组别的扰动强度与监督水平,系统分析了引导强度对模型性能的影响机制。研究发现低置信度样本是导致CO及鲁棒性-准确性权衡的主要因素。基于此洞察,我们提出分布感知动态引导(DDG)策略,该策略可动态调整扰动预算与监督信号。具体而言,DDG根据样本在真实类别上的置信度缩放扰动幅度,从而引导样本趋向一致的决策边界,同时抑制对虚假相关性的学习;并基于每个样本的预测状态动态调整监督信号,避免对错误信号的过度关注。为缓解动态引导可能引发的梯度不稳定性,我们进一步设计了加权正则化约束。在标准基准数据集上的大量实验表明,DDG能有效缓解CO及鲁棒性-准确性权衡问题。