In this paper, we explore the potential of Large Language Models (LLMs) Agents in playing the strategic social deduction game, Resistance Avalon. Players in Avalon are challenged not only to make informed decisions based on dynamically evolving game phases, but also to engage in discussions where they must deceive, deduce, and negotiate with other players. These characteristics make Avalon a compelling test-bed to study the decision-making and language-processing capabilities of LLM Agents. To facilitate research in this line, we introduce AvalonBench - a comprehensive game environment tailored for evaluating multi-agent LLM Agents. This benchmark incorporates: (1) a game environment for Avalon, (2) rule-based bots as baseline opponents, and (3) ReAct-style LLM agents with tailored prompts for each role. Notably, our evaluations based on AvalonBench highlight a clear capability gap. For instance, models like ChatGPT playing good-role got a win rate of 22.2% against rule-based bots playing evil, while good-role bot achieves 38.2% win rate in the same setting. We envision AvalonBench could be a good test-bed for developing more advanced LLMs (with self-playing) and agent frameworks that can effectively model the layered complexities of such game environments.
翻译:本文探索了大型语言模型(LLMs)代理在战略社交推理游戏《抵抗组织:阿瓦隆》中的潜力。阿瓦隆中的玩家不仅需要根据动态演变的游戏阶段做出明智决策,还需参与涉及欺骗、推理和与其他玩家谈判的讨论。这些特性使阿瓦隆成为研究LLM代理决策与语言处理能力的理想测试平台。为促进该方向研究,我们引入AvalonBench——一个专为评估多智能体LLM代理设计的综合性游戏环境。该基准包含:(1)阿瓦隆游戏环境,(2)基于规则的机器人作为基准对手,以及(3)为各角色定制的ReAct风格LLM代理。值得注意的是,基于AvalonBench的评估揭示出明显的性能差距。例如,扮演好角色的ChatGPT在对抗扮演坏角色的规则机器人时获胜率为22.2%,而好角色机器人在相同设置下则获得38.2%的胜率。我们期待AvalonBench能成为开发更先进LLM(通过自我对弈)及能有效建模此类游戏环境分层复杂性的代理框架的优秀测试平台。