Adversarial Attacks are still a significant challenge for neural networks. Recent work has shown that adversarial perturbations typically contain high-frequency features, but the root cause of this phenomenon remains unknown. Inspired by theoretical work on linear full-width convolutional models, we hypothesize that the local (i.e. bounded-width) convolutional operations commonly used in current neural networks are implicitly biased to learn high frequency features, and that this is one of the root causes of high frequency adversarial examples. To test this hypothesis, we analyzed the impact of different choices of linear and nonlinear architectures on the implicit bias of the learned features and the adversarial perturbations, in both spatial and frequency domains. We find that the high-frequency adversarial perturbations are critically dependent on the convolution operation because the spatially-limited nature of local convolutions induces an implicit bias towards high frequency features. The explanation for the latter involves the Fourier Uncertainty Principle: a spatially-limited (local in the space domain) filter cannot also be frequency-limited (local in the frequency domain). Furthermore, using larger convolution kernel sizes or avoiding convolutions (e.g. by using Vision Transformers architecture) significantly reduces this high frequency bias, but not the overall susceptibility to attacks. Looking forward, our work strongly suggests that understanding and controlling the implicit bias of architectures will be essential for achieving adversarial robustness.
翻译:对抗攻击仍是神经网络面临的重大挑战。最新研究表明,对抗扰动通常包含高频特征,但该现象的根本原因尚不明确。受线性全宽度卷积模型理论研究的启发,我们假设当前神经网络中常用的局部(即有界宽度)卷积运算会隐式地偏向学习高频特征,这可能是高频对抗样本产生的根本原因之一。为验证该假设,我们在空间域和频率域中分析了线性和非线性架构的不同选择对学习特征及对抗扰动隐式偏好的影响。研究发现,高频对抗扰动与卷积运算密切相关,因为局部卷积在空间上的有限性会导致对高频特征的隐式偏好。对此现象的解释涉及傅里叶不确定性原理:空间域受限(空间局部)的滤波器无法同时实现频率域受限(频率局部)。此外,采用更大卷积核尺寸或避免使用卷积(如采用视觉Transformer架构)可显著降低这种高频偏好,但不会降低整体攻击敏感性。展望未来,本研究强烈表明,理解并控制架构的隐式偏好对于实现对抗鲁棒性至关重要。