Open-ended learning methods that automatically generate a curriculum of increasingly challenging tasks serve as a promising avenue toward generally capable reinforcement learning agents. Existing methods adapt curricula independently over either environment parameters (in single-agent settings) or co-player policies (in multi-agent settings). However, the strengths and weaknesses of co-players can manifest themselves differently depending on environmental features. It is thus crucial to consider the dependency between the environment and co-player when shaping a curriculum in multi-agent domains. In this work, we use this insight and extend Unsupervised Environment Design (UED) to multi-agent environments. We then introduce Multi-Agent Environment Design Strategist for Open-Ended Learning (MAESTRO), the first multi-agent UED approach for two-player zero-sum settings. MAESTRO efficiently produces adversarial, joint curricula over both environments and co-players and attains minimax-regret guarantees at Nash equilibrium. Our experiments show that MAESTRO outperforms a number of strong baselines on competitive two-player games, spanning discrete and continuous control settings.
翻译:开放式学习方法能够自动生成难度递增的任务课程,为培养具备通用能力的强化学习智能体提供了有前景的途径。现有方法要么独立于环境参数(在单智能体场景中)自适应课程,要么独立于对手策略(在多智能体场景中)进行自适应。然而,对手策略的优势与劣势可能因环境特征的不同而表现出显著差异。因此,在多智能体领域的课程设计中,必须考虑环境与对手策略之间的依赖关系。本研究基于这一洞见,将无监督环境设计(UED)拓展至多智能体环境。我们进一步提出了面向开放式学习的多智能体环境设计策略(MAESTRO),这是首个针对双人零和博弈场景的多智能体UED方法。MAESTRO能够高效地生成同时涵盖环境与对手策略的对抗性联合课程,并在纳什均衡条件下实现极小极大遗憾保证。实验表明,在涵盖离散控制与连续控制场景的竞争性双人博弈中,MAESTRO显著优于多个强基线方法。