Deep equilibrium models (DEQs) refrain from the traditional layer-stacking paradigm and turn to find the fixed point of a single layer. DEQs have achieved promising performance on different applications with featured memory efficiency. At the same time, the adversarial vulnerability of DEQs raises concerns. Several works propose to certify robustness for monotone DEQs. However, limited efforts are devoted to studying empirical robustness for general DEQs. To this end, we observe that an adversarially trained DEQ requires more forward steps to arrive at the equilibrium state, or even violates its fixed-point structure. Besides, the forward and backward tracks of DEQs are misaligned due to the black-box solvers. These facts cause gradient obfuscation when applying the ready-made attacks to evaluate or adversarially train DEQs. Given this, we develop approaches to estimate the intermediate gradients of DEQs and integrate them into the attacking pipelines. Our approaches facilitate fully white-box evaluations and lead to effective adversarial defense for DEQs. Extensive experiments on CIFAR-10 validate the adversarial robustness of DEQs competitive with deep networks of similar sizes.
翻译:深度均衡模型(DEQs)摒弃了传统的层级堆叠范式,转而寻求单层的固定点。DEQs凭借其特有的内存效率,在不同应用场景中取得了令人瞩目的性能。与此同时,DEQs的对抗脆弱性引发了担忧。已有若干工作提出为单调DEQs提供认证鲁棒性,但针对通用DEQs的经验鲁棒性研究仍十分有限。为此,我们观察到:经过对抗训练的DEQ需要更多前向步骤才能达到平衡状态,甚至可能破坏其固定点结构;此外,由于黑盒求解器的存在,DEQ的前向与反向传播轨迹存在错位。这些现象导致在应用现成攻击方法评估或对抗训练DEQ时出现梯度混淆。鉴于此,我们开发了DEQ中间梯度估计方法并将其集成至攻击流程中。我们的方法可实现完全白盒评估,并为DEQ提供有效的对抗防御。在CIFAR-10上的大量实验验证了DEQ的对抗鲁棒性可与同等规模深度网络相媲美。