Large Multimodal Models (LMMs) rely on pre-trained Vision Language Models (VLMs) and Large Language Models (LLMs) to perform amazing emergent abilities on various multimodal tasks in the joint space of vision and language. However, the Typographic Attack, which shows disruption to VLMs, has also been certified as a security vulnerability to LMMs. In this work, we first comprehensively investigate the distractibility of LMMs by typography. In particular, we introduce the Typographic Dataset designed to evaluate distractibility across various multi-modal subtasks, such as object recognition, visual attributes detection, enumeration, arithmetic computation, and commonsense reasoning. To further study the effect of typographic patterns on performance, we also scrutinize the effect of tuning various typographic factors, encompassing font size, color, opacity, and spatial positioning of typos. We discover that LMMs can partially distinguish visual contents and typos when confronting typographic attacks, which suggests that embeddings from vision encoders contain enough information to distinguish visual contents and typos in images. Inspired by such phenomena, we demonstrate that CLIP's performance of zero-shot classification on typo-ridden images can be significantly improved by providing more informative texts to match images. Furthermore, we also prove that LMMs can utilize more informative prompts to leverage information in embeddings to differentiate between visual content and typos. Finally, we propose a prompt information enhancement method that can effectively mitigate the effects of typography.
翻译:大型多模态模型(LMMs)依赖预训练的视觉语言模型(VLMs)和大语言模型(LLMs),在视觉与语言的联合空间中展现出跨多种多模态任务的惊人涌现能力。然而,已被证实能够干扰VLM的印刷攻击,同样被确认为LMMs的安全漏洞。本文首先系统研究了LMMs受印刷文字干扰的可分散性。具体而言,我们引入了专为评估多模态子任务(如物体识别、视觉属性检测、计数、算术计算及常识推理)中可分散性而设计的印刷数据集。为深入探究印刷模式对性能的影响,我们还细致考察了印刷文字的字号、颜色、透明度及空间位置等多类因素调整带来的效果。研究发现,面对印刷攻击时,LMMs能够部分区分视觉内容与印刷文字,这表明视觉编码器生成的嵌入向量已包含足够信息来区分图像中的视觉内容与印刷文字。受此现象启发,我们证明通过提供更丰富文本来匹配图像,能显著提升CLIP在含印刷文字图像上的零样本分类性能。此外,我们还验证了LMMs可通过利用更丰富的提示来提取嵌入向量中的信息,从而区分视觉内容与印刷文字。最终,我们提出一种提示信息增强方法,可有效缓解印刷攻击的影响。