Value function factorization has achieved great success in multi-agent reinforcement learning by optimizing joint action-value functions through the maximization of factorized per-agent utilities. To ensure Individual-Global-Maximum property, existing works often focus on value factorization using monotonic functions, which are known to result in restricted representation expressiveness. In this paper, we analyze the limitations of monotonic factorization and present ConcaveQ, a novel non-monotonic value function factorization approach that goes beyond monotonic mixing functions and employs neural network representations of concave mixing functions. Leveraging the concave property in factorization, an iterative action selection scheme is developed to obtain optimal joint actions during training. It is used to update agents' local policy networks, enabling fully decentralized execution. The effectiveness of the proposed ConcaveQ is validated across scenarios involving multi-agent predator-prey environment and StarCraft II micromanagement tasks. Empirical results exhibit significant improvement of ConcaveQ over state-of-the-art multi-agent reinforcement learning approaches.
翻译:值函数分解通过最大化分解后的每个智能体效用函数来优化联合动作值函数,已在多智能体强化学习中取得巨大成功。为确保个体-全局最大值性质,现有工作通常聚焦于采用单调函数进行值函数分解,但已知这类方法会限制表示能力的表达性。本文分析了单调分解的局限性,并提出ConcaveQ——一种突破单调混合函数的新型非单调值函数分解方法,该方法采用凹混合函数的神经网络表示。利用分解中的凹性性质,本文开发了一种迭代动作选择方案,在训练过程中获取最优联合动作。该方案用于更新智能体的局部策略网络,从而实现完全去中心化的执行。通过多智能体捕食者-猎物环境与星际争霸II微观管理任务场景验证了所提ConcaveQ的有效性。实验结果表明,ConcaveQ相较于现有最先进的多智能体强化学习方法实现了显著性能提升。