Recent breakthroughs in large language models (LLMs) have brought remarkable success in the field of LLM-as-Agent. Nevertheless, a prevalent assumption is that the information processed by LLMs is consistently honest, neglecting the pervasive deceptive or misleading information in human society and AI-generated content. This oversight makes LLMs susceptible to malicious manipulations, potentially resulting in detrimental outcomes. This study utilizes the intricate Avalon game as a testbed to explore LLMs' potential in deceptive environments. Avalon, full of misinformation and requiring sophisticated logic, manifests as a "Game-of-Thoughts". Inspired by the efficacy of humans' recursive thinking and perspective-taking in the Avalon game, we introduce a novel framework, Recursive Contemplation (ReCon), to enhance LLMs' ability to identify and counteract deceptive information. ReCon combines formulation and refinement contemplation processes; formulation contemplation produces initial thoughts and speech, while refinement contemplation further polishes them. Additionally, we incorporate first-order and second-order perspective transitions into these processes respectively. Specifically, the first-order allows an LLM agent to infer others' mental states, and the second-order involves understanding how others perceive the agent's mental state. After integrating ReCon with different LLMs, extensive experiment results from the Avalon game indicate its efficacy in aiding LLMs to discern and maneuver around deceptive information without extra fine-tuning and data. Finally, we offer a possible explanation for the efficacy of ReCon and explore the current limitations of LLMs in terms of safety, reasoning, speaking style, and format, potentially furnishing insights for subsequent research.
翻译:大型语言模型(LLM)在LLM-as-Agent领域取得了突破性进展。然而,一个普遍存在的假设是,LLM处理的信息始终是诚实的,忽视了人类社会和人工智能生成内容中普遍存在的欺骗性或误导性信息。这种疏忽使LLM容易受到恶意操纵,可能导致有害后果。本研究利用复杂的阿瓦隆游戏作为测试平台,探索LLM在欺骗性环境中的潜力。充满错误信息且需要复杂逻辑推理的阿瓦隆游戏,本质上呈现为一场"思维博弈"。受人类在阿瓦隆游戏中递归思维和换位思考有效性的启发,我们提出了一种新颖框架——递归沉思(ReCon),以增强LLM识别和对抗欺骗性信息的能力。ReCon结合了构思与精炼两个沉思过程:构思过程产生初始想法和言语,而精炼过程则进一步优化它们。此外,我们分别将一阶和二阶视角转换融入这些过程。具体而言,一阶视角转换允许LLM智能体推断他人的心理状态,二阶视角转换则涉及理解他人如何看待智能体自身的心理状态。将ReCon与不同LLM集成后,阿瓦隆游戏的大量实验结果表明,该框架无需额外微调和数据即可有效帮助LLM识别并应对欺骗性信息。最后,我们为ReCon的有效性提供了可能的解释,并探讨了当前LLM在安全性、推理、说话风格和格式方面的局限性,这有望为后续研究提供启示。