Conventions are crucial for strong performance in cooperative multi-agent games, because they allow players to coordinate on a shared strategy without explicit communication. Unfortunately, standard multi-agent reinforcement learning techniques, such as self-play, converge to conventions that are arbitrary and non-diverse, leading to poor generalization when interacting with new partners. In this work, we present a technique for generating diverse conventions by (1) maximizing their rewards during self-play, while (2) minimizing their rewards when playing with previously discovered conventions (cross-play), stimulating conventions to be semantically different. To ensure that learned policies act in good faith despite the adversarial optimization of cross-play, we introduce \emph{mixed-play}, where an initial state is randomly generated by sampling self-play and cross-play transitions and the player learns to maximize the self-play reward from this initial state. We analyze the benefits of our technique on various multi-agent collaborative games, including Overcooked, and find that our technique can adapt to the conventions of humans, surpassing human-level performance when paired with real users.
翻译:惯例对于合作型多智能体游戏中的优异表现至关重要,因为它们使玩家无需明确通信即可协调共享策略。然而,标准的多智能体强化学习技术(如自我对弈)会收敛到任意且非多样化的惯例,导致在与新伙伴交互时泛化能力较差。本研究提出一种生成多样化惯例的技术,通过(1)在自我对弈中最大化其奖励,同时(2)在与先前发现的惯例(交叉对弈)对弈时最小化其奖励,激励惯例在语义上产生差异。为确保学习策略在交叉对弈的对抗性优化中仍保持善意,我们引入混合对弈机制:通过采样自我对弈和交叉对弈的转移过程随机生成初始状态,智能体从该初始状态学习最大化自我对弈奖励。我们分析了该技术在Overcooked等多人协作游戏中的优势,发现该技术能适应人类惯例,在与真实用户配对时超越人类水平表现。