Multi-Agent Reinforcement Learning (MARL) is a promising area of research that can model and control multiple, autonomous decision-making agents. During online training, MARL algorithms involve performance-intensive computations such as exploration and exploitation phases originating from large observation-action space belonging to multiple agents. In this article, we seek to characterize the scalability bottlenecks in several popular classes of MARL algorithms during their training phases. Our experimental results reveal new insights into the key modules of MARL algorithms that limit the scalability, and outline potential strategies that may help address these performance issues.
翻译:多智能体强化学习(Multi-Agent Reinforcement Learning, MARL)是一个前景广阔的研究领域,能够对多个自主决策智能体进行建模与控制。在在线训练过程中,多智能体强化学习算法需执行高强度的计算任务,例如源于多个智能体的大规模观测-动作空间的探索与利用阶段。本文旨在刻画几类主流多智能体强化学习算法在训练阶段的可扩展性瓶颈。我们的实验结果揭示了限制可扩展性的多智能体强化学习算法关键模块的新见解,并概述了可能有助于应对这些性能问题的潜在策略。