Multi-agent reinforcement learning (MARL) has witnessed significant progress with the development of value function factorization methods. It allows optimizing a joint action-value function through the maximization of factorized per-agent utilities due to monotonicity. In this paper, we show that in partially observable MARL problems, an agent's ordering over its own actions could impose concurrent constraints (across different states) on the representable function class, causing significant estimation error during training. We tackle this limitation and propose PAC, a new framework leveraging Assistive information generated from Counterfactual Predictions of optimal joint action selection, which enable explicit assistance to value function factorization through a novel counterfactual loss. A variational inference-based information encoding method is developed to collect and encode the counterfactual predictions from an estimated baseline. To enable decentralized execution, we also derive factorized per-agent policies inspired by a maximum-entropy MARL framework. We evaluate the proposed PAC on multi-agent predator-prey and a set of StarCraft II micromanagement tasks. Empirical results demonstrate improved results of PAC over state-of-the-art value-based and policy-based multi-agent reinforcement learning algorithms on all benchmarks.
翻译:多智能体强化学习(MARL)随着值函数分解方法的发展取得了显著进步。该方法通过最大化因式分解的每个智能体效用(由于单调性)来优化联合动作值函数。本文证明,在部分可观测的MARL问题中,智能体对其自身动作的排序可能在不同状态下对可表示函数类施加并发约束,导致训练过程中产生显著估计误差。为解决此局限性,我们提出PAC框架,该框架利用从最优联合动作选择的反事实预测中生成的辅助信息,通过新颖的反事实损失对值函数分解提供显式辅助。我们开发了基于变分推理的信息编码方法,用于收集并编码来自估计基线的反事实预测。为实现去中心化执行,我们受最大熵MARL框架启发推导了因式分解的每个智能体策略。在多个智能体捕食者-猎物任务及一组星际争霸II微观管理任务上评估了所提出的PAC方法。实证结果表明,在所有基准测试中,PAC相较于最先进的基于值和基于策略的多智能体强化学习算法均取得了更优结果。