Vision-language models (VLMs) seamlessly integrate visual and textual data to perform tasks such as image classification, caption generation, and visual question answering. However, adversarial images often struggle to deceive all prompts effectively in the context of cross-prompt migration attacks, as the probability distribution of the tokens in these images tends to favor the semantics of the original image rather than the target tokens. To address this challenge, we propose a Contextual-Injection Attack (CIA) that employs gradient-based perturbation to inject target tokens into both visual and textual contexts, thereby improving the probability distribution of the target tokens. By shifting the contextual semantics towards the target tokens instead of the original image semantics, CIA enhances the cross-prompt transferability of adversarial images.Extensive experiments on the BLIP2, InstructBLIP, and LLaVA models show that CIA outperforms existing methods in cross-prompt transferability, demonstrating its potential for more effective adversarial strategies in VLMs.
翻译:视觉语言模型(VLMs)无缝整合视觉与文本数据,以执行图像分类、描述生成和视觉问答等任务。然而,在跨提示迁移攻击的背景下,对抗性图像往往难以有效欺骗所有提示,因为这些图像中标记的概率分布倾向于偏向原始图像的语义而非目标标记。为应对这一挑战,我们提出一种上下文注入攻击(CIA),该方法采用基于梯度的扰动将目标标记注入视觉和文本上下文中,从而改善目标标记的概率分布。通过将上下文语义向目标标记而非原始图像语义转移,CIA增强了对抗性图像的跨提示可迁移性。在BLIP2、InstructBLIP和LLaVA模型上的大量实验表明,CIA在跨提示可迁移性方面优于现有方法,证明了其在视觉语言模型中实现更有效对抗策略的潜力。