Understanding how humans collaborate and communicate in teams is essential for improving human-agent teaming and AI-assisted decision-making. However, relying solely on data from large-scale user studies is impractical due to logistical, ethical, and practical constraints, necessitating synthetic models of multiple diverse human behaviors. Recently, agents powered by Large Language Models (LLMs) have been shown to emulate human-like behavior in social settings. But, obtaining a large set of diverse behaviors requires manual effort in the form of designing prompts. On the other hand, Quality Diversity (QD) optimization has been shown to be capable of generating diverse Reinforcement Learning (RL) agent behavior. In this work, we combine QD optimization with LLM-powered agents to iteratively search for prompts that generate diverse team behavior in a long-horizon, multi-step collaborative environment. We first show, through a human-subjects experiment, that humans exhibit diverse coordination and communication behavior in this domain. We then present a series of experiments showing that our approach captures behaviors that are difficult to observe without large-scale data collection, and a follow-up user study to show that these generated behaviors are human-like. Our findings highlight the combination of QD and LLM-powered agents as an effective tool for studying teaming and communication strategies in multi-agent collaboration.
翻译:理解人类如何在团队中协作与通信对于提升人机协同及AI辅助决策至关重要。然而,仅依赖大规模用户研究数据因后勤、伦理及实践限制而不可行,亟需合成多类多样化人类行为的模型。近期研究表明,基于大语言模型的智能体能够在社交情境中模拟类人行为。但获取大量多样化行为需通过设计提示进行人工操作。另一方面,质量多样性优化已被证实在强化学习智能体行为生成中具有多样性。本研究将质量多样性优化与大语言模型驱动的智能体相结合,在长时域多步协作环境中迭代搜索生成多样化团队行为的提示。我们首先通过人类受试者实验证明,人类在此领域展现出多样化的协调与通信行为。随后通过系列实验表明,我们的方法能够捕捉到缺乏大规模数据收集时难以观测的行为,并开展后续用户研究证实这些生成的行为具有类人性。研究结果凸显了质量多样性优化与大语言模型智能体的结合,是研究多智能体协作中团队协作与通信策略的有效工具。