The rapid growth and increasing popularity of incorporating additional modalities (e.g., vision) into large language models (LLMs) has raised significant security concerns. This expansion of modality, akin to adding more doors to a house, unintentionally creates multiple access points for adversarial attacks. In this paper, by introducing adversarial embedding space attacks, we emphasize the vulnerabilities present in multi-modal systems that originate from incorporating off-the-shelf components like public pre-trained encoders in a plug-and-play manner into these systems. In contrast to existing work, our approach does not require access to the multi-modal system's weights or parameters but instead relies on the huge under-explored embedding space of such pre-trained encoders. Our proposed embedding space attacks involve seeking input images that reside within the dangerous or targeted regions of the extensive embedding space of these pre-trained components. These crafted adversarial images pose two major threats: 'Context Contamination' and 'Hidden Prompt Injection'-both of which can compromise multi-modal models like LLaVA and fully change the behavior of the associated language model. Our findings emphasize the need for a comprehensive examination of the underlying components, particularly pre-trained encoders, before incorporating them into systems in a plug-and-play manner to ensure robust security.
翻译:大型语言模型(LLMs)中快速整合额外模态(如视觉)的趋势日益增长,引发了显著的安全担忧。这种模态扩展如同为房屋增加更多门户,无意中为对抗性攻击创造了多个入口点。本文通过引入对抗性嵌入空间攻击,揭示了多模态系统中因采用即插即用方式整合现成组件(如公开预训练编码器)所固有的脆弱性。与现有研究不同,我们的方法无需访问多模态系统的权重或参数,而是依赖于这类预训练编码器未被充分探索的广阔嵌入空间。我们提出的嵌入空间攻击旨在寻找位于这些预训练组件庞大嵌入空间中危险或目标区域的输入图像。这些精心构造的对抗性图像构成两大威胁:"上下文污染"和"隐藏提示注入"——两者均可破坏LLaVA等多模态模型,并彻底改变关联语言模型的行为。我们的研究强调,在以即插即用方式将底层组件(尤其是预训练编码器)整合到系统前,必须对其进行全面审查,以确保鲁棒的安全性。