In many coordination problems, independently reasoning humans are able to discover mutually compatible policies. In contrast, independently trained self-play policies are often mutually incompatible. Zero-shot coordination (ZSC) has recently been proposed as a new frontier in multi-agent reinforcement learning to address this fundamental issue. Prior work approaches the ZSC problem by assuming players can agree on a shared learning algorithm but not on labels for actions and observations, and proposes other-play as an optimal solution. However, until now, this "label-free" problem has only been informally defined. We formalize this setting as the label-free coordination (LFC) problem by defining the label-free coordination game. We show that other-play is not an optimal solution to the LFC problem as it fails to consistently break ties between incompatible maximizers of the other-play objective. We introduce an extension of the algorithm, other-play with tie-breaking, and prove that it is optimal in the LFC problem and an equilibrium in the LFC game. Since arbitrary tie-breaking is precisely what the ZSC setting aims to prevent, we conclude that the LFC problem does not reflect the aims of ZSC. To address this, we introduce an alternative informal operationalization of ZSC as a starting point for future work.
翻译:在许多协调问题中,独立推理的人类能够发现相互兼容的策略。相比之下,独立训练的自我博弈策略往往相互不兼容。零样本协调(ZSC)最近被提出作为多智能体强化学习中的一个新前沿,以解决这一根本性问题。先前的工作通过假设玩家能够就共享的学习算法达成一致,但无法就动作和观测的标签达成一致来研究ZSC问题,并提出其他博弈(other-play)作为最优解。然而,迄今为止,这一“无标签”问题仅被非正式地定义。我们通过定义无标签协调博弈,将这一场景形式化为无标签协调(LFC)问题。我们证明其他博弈并非LFC问题的最优解,因为它无法一致地打破其他博弈目标函数中不兼容的最大化者之间的平局。我们引入了一种算法扩展——带打破平局的其他博弈,并证明它在LFC问题中是最优的,且在LFC博弈中是一个均衡。由于任意打破平局正是ZSC场景旨在避免的,我们得出结论:LFC问题并未反映ZSC的目标。为此,我们引入了一种替代的非正式ZSC操作性定义,作为未来工作的起点。