Leveraging the rapid development of Large Language Models LLMs, LLM-based agents have been developed to handle various real-world applications, including finance, healthcare, and shopping, etc. It is crucial to ensure the reliability and security of LLM-based agents during applications. However, the safety issues of LLM-based agents are currently under-explored. In this work, we take the first step to investigate one of the typical safety threats, backdoor attack, to LLM-based agents. We first formulate a general framework of agent backdoor attacks, then we present a thorough analysis on the different forms of agent backdoor attacks. Specifically, from the perspective of the final attacking outcomes, the attacker can either choose to manipulate the final output distribution, or only introduce malicious behavior in the intermediate reasoning process, while keeping the final output correct. Furthermore, the former category can be divided into two subcategories based on trigger locations: the backdoor trigger can be hidden either in the user query or in an intermediate observation returned by the external environment. We propose the corresponding data poisoning mechanisms to implement the above variations of agent backdoor attacks on two typical agent tasks, web shopping and tool utilization. Extensive experiments show that LLM-based agents suffer severely from backdoor attacks, indicating an urgent need for further research on the development of defenses against backdoor attacks on LLM-based agents. Warning: This paper may contain biased content.
翻译:随着大语言模型(LLM)的快速发展,基于LLM的智能体被开发出来处理各类实际应用,涵盖金融、医疗、购物等领域。在应用过程中,确保基于LLM的智能体的可靠性和安全性至关重要。然而,目前针对此类智能体的安全问题研究尚不充分。本研究首次系统探究了基于LLM的智能体面临的一类典型安全威胁——后门攻击。我们首先构建了智能体后门攻击的通用框架,随后对后门攻击的不同形式进行了深入分析。具体而言,从最终攻击效果的角度看,攻击者既可以操纵最终输出分布,也可以在中间推理过程中引入恶意行为而保持最终输出正确。此外,前一类攻击可根据触发位置进一步分为两个子类:后门触发器可隐藏于用户查询中,或隐藏在外部环境返回的中间观察结果里。我们提出了相应的数据投毒机制,在两个典型智能体任务(网络购物与工具使用)上实现了上述后门攻击变体。大量实验表明,基于LLM的智能体深受后门攻击影响,亟需开展针对此类智能体后门攻击防御机制的进一步研究。警告:本论文可能包含偏见内容。