Progress in fields of machine learning and adversarial planning has benefited significantly from benchmark domains, from checkers and the classic UCI data sets to Go and Diplomacy. In sequential decision-making, agent evaluation has largely been restricted to few interactions against experts, with the aim to reach some desired level of performance (e.g. beating a human professional player). We propose a benchmark for multiagent learning based on repeated play of the simple game Rock, Paper, Scissors along with a population of forty-three tournament entries, some of which are intentionally sub-optimal. We describe metrics to measure the quality of agents based both on average returns and exploitability. We then show that several RL, online learning, and language model approaches can learn good counter-strategies and generalize well, but ultimately lose to the top-performing bots, creating an opportunity for research in multiagent learning.
翻译:机器学习和对抗规划领域的进展很大程度上受益于基准领域,从跳棋和经典UCI数据集到围棋和外交。在序列决策中,智能体评估主要局限于与专家进行少量交互,目标是达到某种期望的性能水平(例如,击败人类职业选手)。我们提出一个基于简单游戏“石头、剪刀、布”重复对局的多智能体学习基准,并包含一个由四十三个锦标赛参赛者组成的种群,其中一些有意设计为次优。我们描述了基于平均收益和可剥削性来测量智能体质量的指标。然后,我们展示了若干强化学习、在线学习和语言模型方法能够学习良好的对抗策略并良好地泛化,但最终仍输给性能最高的机器人,这为多智能体学习的研究创造了机会。