Asymmetrical multiplayer (AMP) game is a popular game genre which involves multiple types of agents competing or collaborating with each other in the game. It is difficult to train powerful agents that can defeat top human players in AMP games by typical self-play training method because of unbalancing characteristics in their asymmetrical environments. We propose asymmetric-evolution training (AET), a novel multi-agent reinforcement learning framework that can train multiple kinds of agents simultaneously in AMP game. We designed adaptive data adjustment (ADA) and environment randomization (ER) to optimize the AET process. We tested our method in a complex AMP game named Tom \& Jerry, and our AIs trained without using any human data can achieve a win rate of 98.5% against top human players over 65 matches. The ablation experiments indicated that the proposed modules are beneficial to the framework.
翻译:非对称多人游戏是一种流行的游戏类型,涉及多种类型的智能体在游戏中相互竞争或合作。由于非对称环境中的不平衡特性,使用典型的自我对弈训练方法难以训练出能够击败顶尖人类玩家的强大智能体。我们提出了非对称演化训练(AET),这是一种新颖的多智能体强化学习框架,能够在非对称多人游戏中同时训练多种类型的智能体。我们设计了自适应数据调整(ADA)和环境随机化(ER)来优化AET过程。我们在一个名为《汤姆与杰瑞》的复杂非对称多人游戏中测试了该方法,我们的AI在未使用任何人机数据的情况下,与顶尖人类玩家进行了65场比赛,胜率达到98.5%。消融实验表明,所提出的模块对框架具有正面作用。