While LLM-based agents excel at individual tasks, effective collaboration with realistic human partners remains challenging. Most of the existing conversation-level collaborative studies lack grounded interaction and behavioral execution, motivating the need for cooperative game environments that enable contextualized and immersive collaboration. To this end, this paper proposes CollabBench, a benchmark for evaluating and training collaborative agents in cooperative games. CollabBench features a Diverse Player Profile Simulation pipeline to model varied players behaviors, and a Collaborative Agentic Training paradigm that unifies reasoning, communication, and action via agentic rollouts, optimized with a hybrid reward balancing task efficiency and affective adaptation. We further extend classic environments to CWAH-MultiPlayer and Cook-MultiPlayer for systematic evaluation under diverse personalities. Experiments with efficiency and affective metrics show that our trained models outperform base models, achieving 19.5% higher efficiency and 24.4% improved affective performance. Further analysis reveals key collaborative limitations of existing models and offers insights for future collaborative training.
翻译:基于大语言模型的智能体在独立任务中表现出色,但实现与真实人类伙伴的有效协作仍具挑战性。现有的大多数对话级协作研究缺乏具身交互与行为执行,这促使我们需要能实现情境化沉浸式协作的合作游戏环境。为此,本文提出CollabBench——面向合作游戏的协作智能体评估与训练基准。CollabBench包含多样玩家画像模拟流程,用以建模不同玩家的行为模式,以及统一推理、通信与动作的协作智能体训练范式,该范式通过智能体式推演实现协同优化,并采用混合奖励机制平衡任务效率与情感适应性。我们进一步将经典环境扩展为CWAH-MultiPlayer与Cook-MultiPlayer,以支持不同人格特征下的系统性评估。基于效率与情感指标的实验表明,训练后的模型相较基础模型实现19.5%的效率提升与24.4%的情感表现改进。进一步分析揭示了现有模型的关键协作局限,并为未来协作训练提供启示。