Adversarial training has proven effective in improving the robustness of deep neural networks against adversarial attacks. However, this enhanced robustness often comes at the cost of a substantial drop in accuracy on clean data. In this paper, we address this limitation by introducing Tangent Direction Guided Adversarial Training (TART), a novel method that enhances clean accuracy by exploiting the geometry of the data manifold. We argue that adversarial examples with large components in the normal direction can overly distort the decision boundary and degrade clean accuracy. TART addresses this issue by estimating the tangent direction of adversarial examples and adaptively modulating the perturbation bound based on the norm of their tangential component. To the best of our knowledge, TART is the first adversarial defense framework that explicitly incorporates the concept of tangent space and direction into adversarial training. Extensive experiments on both synthetic and benchmark datasets demonstrate that TART consistently improves clean accuracy while maintaining robustness against adversarial attacks.
翻译:对抗训练已被证明能有效提升深度神经网络对对抗攻击的鲁棒性。然而,这种增强的鲁棒性通常以在干净数据上精度大幅下降为代价。在本文中,我们通过引入切线方向引导对抗训练(TART)这一新颖方法来解决这一限制。该方法利用数据流形的几何特性来提升干净精度。我们认为,在法线方向具有较大分量的对抗样本会过度扭曲决策边界,从而降低干净精度。TART通过估计对抗样本的切线方向,并根据其切向分量的范数自适应调整扰动边界来解决这一问题。据我们所知,TART是首个将切空间和切线方向概念显式融入对抗训练的防御框架。在合成数据集和基准数据集上的大量实验表明,TART在维持对抗攻击鲁棒性的同时,能持续提升干净精度。