Pretrained vision-language models (VLMs) like CLIP have shown impressive generalization performance across various downstream tasks, yet they remain vulnerable to adversarial attacks. While prior research has primarily concentrated on improving the adversarial robustness of image encoders to guard against attacks on images, the exploration of text-based and multimodal attacks has largely been overlooked. In this work, we initiate the first known and comprehensive effort to study adapting vision-language models for adversarial robustness under the multimodal attack. Firstly, we introduce a multimodal attack strategy and investigate the impact of different attacks. We then propose a multimodal contrastive adversarial training loss, aligning the clean and adversarial text embeddings with the adversarial and clean visual features, to enhance the adversarial robustness of both image and text encoders of CLIP. Extensive experiments on 15 datasets across two tasks demonstrate that our method significantly improves the adversarial robustness of CLIP. Interestingly, we find that the model fine-tuned against multimodal adversarial attacks exhibits greater robustness than its counterpart fine-tuned solely against image-based attacks, even in the context of image attacks, which may open up new possibilities for enhancing the security of VLMs.
翻译:预训练的视觉语言模型(如CLIP)在下游任务中展现出令人印象深刻的泛化性能,但面对对抗攻击仍存在脆弱性。尽管先前研究主要集中于提升图像编码器的对抗鲁棒性以防御针对图像的攻击,但基于文本和多模态攻击的探索在很大程度上被忽视。本文首次开展全面研究,旨在探索在**多模态攻击**下如何适应视觉语言模型以实现对抗鲁棒性。首先,我们提出一种多模态攻击策略,并考察不同攻击的影响。随后,我们设计一种多模态对比对抗训练损失,通过对齐干净文本与对抗文本的嵌入表示及对抗视觉特征与干净视觉特征,增强CLIP图像与文本编码器的对抗鲁棒性。在15个数据集上的跨两项任务的大量实验表明,我们的方法显著提升了CLIP的对抗鲁棒性。有趣的是,我们发现在图像攻击场景中,针对多模态对抗攻击微调的模型,其鲁棒性甚至优于仅针对图像攻击微调的模型。这一发现可能为增强视觉语言模型的安全性开辟新途径。