Recent advances in vision-language models (VLMs) have significantly enhanced the visual grounding task, which involves locating objects in an image based on natural language queries. Despite these advancements, the security of VLM-based grounding systems has not been thoroughly investigated. This paper reveals a novel and realistic vulnerability: the first multi-target backdoor attack on VLM-based visual grounding. Unlike prior attacks that rely on static triggers or fixed targets, we propose IAG, a method that dynamically generates input-aware, text-guided triggers conditioned on any specified target object description to execute the attack. This is achieved through a text-conditioned UNet that embeds imperceptible target semantic cues into visual inputs while preserving normal grounding performance on benign samples. We further develop a joint training objective that balances language capability with perceptual reconstruction to ensure imperceptibility, effectiveness, and stealth. Extensive experiments on multiple VLMs (e.g., LLaVA, InternVL, Ferret) and benchmarks (RefCOCO, RefCOCO+, RefCOCOg, Flickr30k Entities, and ShowUI) demonstrate that IAG achieves the best ASRs compared with other baselines on almost all settings without compromising clean accuracy, maintaining robustness against existing defenses, and exhibiting transferability across datasets and models. These findings underscore critical security risks in grounding-capable VLMs and highlight the need for further research on trustworthy multimodal understanding.
翻译:近期视觉语言模型(VLMs)的进展显著提升了视觉定位任务(即根据自然语言查询在图像中定位目标物体)的性能。然而,基于VLM的定位系统安全性尚未得到充分探究。本文揭示了一种新颖且具有现实威胁的漏洞:首个针对VLM视觉定位的多目标后门攻击方法。与依赖静态触发器或固定目标的传统攻击不同,我们提出IAG方法,该方法能动态生成输入感知、文本引导的触发器,并可根据任意指定目标物体描述执行攻击。该方案通过文本条件化UNet在良性样本中嵌入不可感知的目标语义线索至视觉输入,同时保持正常定位性能。我们进一步设计联合训练目标函数,平衡语言能力与感知重建,确保攻击的隐蔽性、有效性和不可察觉性。在多个VLM模型(如LLaVA、InternVL、Ferret)及基准数据集(RefCOCO、RefCOCO+、RefCOCOg、Flickr30k Entities、ShowUI)上的大量实验表明:IAG在几乎所有场景下均实现最优攻击成功率(ASR),且不影响干净样本精度,对现有防御手段具有鲁棒性,并展现出跨数据集和模型的迁移能力。这些发现凸显了具备定位能力的VLM中的关键安全风险,亟需进一步研究可信多模态理解机制。