Deep learning-based Natural Language Processing (NLP) models are vulnerable to adversarial attacks, where small perturbations can cause a model to misclassify. Adversarial Training (AT) is often used to increase model robustness. However, we have discovered an intriguing phenomenon: deliberately or accidentally miscalibrating models masks gradients in a way that interferes with adversarial attack search methods, giving rise to an apparent increase in robustness. We show that this observed gain in robustness is an illusion of robustness (IOR), and demonstrate how an adversary can perform various forms of test-time temperature calibration to nullify the aforementioned interference and allow the adversarial attack to find adversarial examples. Hence, we urge the NLP community to incorporate test-time temperature scaling into their robustness evaluations to ensure that any observed gains are genuine. Finally, we show how the temperature can be scaled during \textit{training} to improve genuine robustness.
翻译:基于深度学习的自然语言处理(NLP)模型易受对抗攻击,即微小扰动可导致模型误分类。对抗训练(AT)常被用于提高模型鲁棒性。然而,我们发现了一个有趣的现象:有意或无意地误校准模型会以干扰对抗攻击搜索方法的方式掩盖梯度,从而产生鲁棒性增强的假象。我们指出这种观察到的鲁棒性提升是一种鲁棒性的幻象(IOR),并展示了攻击者如何通过执行多种形式的测试时温度校准来消除上述干扰,从而使对抗攻击能够找到对抗样本。因此,我们敦促NLP社区将测试时温度缩放纳入鲁棒性评估,以确保任何观察到的提升是真实的。最后,我们展示了如何在训练期间缩放温度以提升真正的鲁棒性。