Due to their unprecedented ability to process and respond to various types of data, Multimodal Large Language Models (MLLMs) are constantly defining the new boundary of Artificial General Intelligence (AGI). As these advanced generative models increasingly form collaborative networks for complex tasks, the integrity and security of these systems are crucial. Our paper, ``The Wolf Within'', explores a novel vulnerability in MLLM societies - the indirect propagation of malicious content. Unlike direct harmful output generation for MLLMs, our research demonstrates how a single MLLM agent can be subtly influenced to generate prompts that, in turn, induce other MLLM agents in the society to output malicious content. This subtle, yet potent method of indirect influence marks a significant escalation in the security risks associated with MLLMs. Our findings reveal that, with minimal or even no access to MLLMs' parameters, an MLLM agent, when manipulated to produce specific prompts or instructions, can effectively ``infect'' other agents within a society of MLLMs. This infection leads to the generation and circulation of harmful outputs, such as dangerous instructions or misinformation, across the society. We also show the transferability of these indirectly generated prompts, highlighting their possibility in propagating malice through inter-agent communication. This research provides a critical insight into a new dimension of threat posed by MLLMs, where a single agent can act as a catalyst for widespread malevolent influence. Our work underscores the urgent need for developing robust mechanisms to detect and mitigate such covert manipulations within MLLM societies, ensuring their safe and ethical utilization in societal applications. Our implementation is released at \url{https://github.com/ChengshuaiZhao0/The-Wolf-Within.git}.
翻译:多模态大语言模型(MLLMs)凭借其处理并响应多种数据类型的空前能力,正持续定义通用人工智能(AGI)的新边界。随着这些先进生成模型日益形成协作网络以处理复杂任务,其系统的完整性与安全性至关重要。我们的论文《狼心潜伏》揭示了MLLM社群中的一种新型脆弱性——恶意内容的间接传播。不同于MLLM直接生成有害输出的方式,本研究展示了如何通过微调诱导单个MLLM智能体生成提示词,进而促使社群中其他MLLM智能体输出恶意内容。这种隐蔽而强大的间接影响方式,标志着与MLLM相关的安全风险出现显著升级。研究发现,即便在极少甚至完全无法获取MLLM参数的情况下,当单个MLLM智能体被操控生成特定提示或指令时,可有效"感染"MLLM社群中的其他智能体。这种感染会导致有害输出(如危险指令或虚假信息)在社群内生成并传播。我们还展示了这些间接生成提示词的可迁移性,揭示了其通过智能体间通信传播恶意内容的可能性。本研究为MLLM构成的新型威胁维度提供了关键洞察:单个智能体即可充当广泛恶意影响的催化剂。我们的工作凸显了开发稳健机制来检测并缓解MLLM社群中此类隐匿操纵的迫切性,以确保其在社会应用中安全且合乎伦理地部署。相关实现已发布于\url{https://github.com/ChengshuaiZhao0/The-Wolf-Within.git}。