The rapid advancements in large language models (LLMs) have led to a resurgence in LLM-based agents, which demonstrate impressive human-like behaviors and cooperative capabilities in various interactions and strategy formulations. However, evaluating the safety of LLM-based agents remains a complex challenge. This paper elaborately conducts a series of manual jailbreak prompts along with a virtual chat-powered evil plan development team, dubbed Evil Geniuses, to thoroughly probe the safety aspects of these agents. Our investigation reveals three notable phenomena: 1) LLM-based agents exhibit reduced robustness against malicious attacks. 2) the attacked agents could provide more nuanced responses. 3) the detection of the produced improper responses is more challenging. These insights prompt us to question the effectiveness of LLM-based attacks on agents, highlighting vulnerabilities at various levels and within different role specializations within the system/agent of LLM-based agents. Extensive evaluation and discussion reveal that LLM-based agents face significant challenges in safety and yield insights for future research. Our code is available at https://github.com/T1aNS1R/Evil-Geniuses.
翻译:大语言模型(LLM)的快速发展带来了基于LLM的智能体复兴,这些智能体在各种交互和策略制定中展现出令人印象深刻的人类行为与合作能力。然而,评估基于LLM智能体的安全性仍然是一个复杂挑战。本文精心设计了一系列手工越狱提示,并构建了一个名为"邪恶天才"的虚拟聊天驱动的邪恶计划开发团队,以深入探测这些智能体的安全方面。我们的研究揭示了三个显著现象:1)基于LLM的智能体对恶意攻击的鲁棒性降低;2)受攻击的智能体可能提供更细微的响应;3)对生成的不当响应的检测更具挑战性。这些见解促使我们质疑基于LLM的智能体攻击的有效性,揭示了系统/智能体内部不同层级及不同角色专业化中的脆弱性。广泛的评估与讨论表明,基于LLM的智能体在安全性方面面临重大挑战,并为未来研究提供了启示。我们的代码可在 https://github.com/T1aNS1R/Evil-Geniuses 获取。