Pretrained large-scale vision-language models like CLIP have exhibited strong generalization over unseen tasks. Yet imperceptible adversarial perturbations can significantly reduce CLIP's performance on new tasks. In this work, we identify and explore the problem of \emph{adapting large-scale models for zero-shot adversarial robustness}. We first identify two key factors during model adaption -- training losses and adaptation methods -- that affect the model's zero-shot adversarial robustness. We then propose a text-guided contrastive adversarial training loss, which aligns the text embeddings and the adversarial visual features with contrastive learning on a small set of training data. We apply this training loss to two adaption methods, model finetuning and visual prompt tuning. We find that visual prompt tuning is more effective in the absence of texts, while finetuning wins in the existence of text guidance. Overall, our approach significantly improves the zero-shot adversarial robustness over CLIP, seeing an average improvement of over 31 points over ImageNet and 15 zero-shot datasets. We hope this work can shed light on understanding the zero-shot adversarial robustness of large-scale models.
翻译:预训练的大规模视觉-语言模型(如CLIP)在未见任务上展现出强大的泛化能力。然而,难以察觉的对抗扰动会显著降低CLIP在新任务上的性能。本文识别并探索了“使大规模模型适应零样本对抗鲁棒性”这一难题。我们首先确定了模型适应过程中的两个关键因素——训练损失和适应方法——它们会影响模型的零样本对抗鲁棒性。随后,我们提出了一种文本引导的对比对抗训练损失,该损失通过对比学习在少量训练数据上对齐文本嵌入与对抗性视觉特征。我们将此训练损失应用于两种适应方法:模型微调和视觉提示调优。我们发现,在缺乏文本引导时,视觉提示调优更为有效;而在存在文本引导时,微调则表现更优。总体而言,我们的方法显著提升了CLIP的零样本对抗鲁棒性,在ImageNet和15个零样本数据集上平均提升了超过31个百分点。我们希望这项工作能为理解大规模模型的零样本对抗鲁棒性提供启示。