Offline cooperative multi-agent reinforcement learning (MARL) faces unique challenges due to distributional shifts, particularly stemming from the high dimensionality of joint action spaces and the presence of out-of-distribution joint action selections. In this work, we highlight that a fundamental challenge in offline MARL arises from the multi-equilibrium nature of cooperative tasks, which induces a highly multimodal joint behavior policy space coupled with heterogeneous-quality behavior data. This makes it difficult for individual policy regularization to align with a consistent coordination pattern, leading to the policy distribution shift problems. To tackle this challenge, we design a sequential score function decomposition method that distills per-agent regularization signals from the joint behavior policy, which induces coordinated modality selection under decentralized execution constraints. Then we leverage a flexible diffusion-based generative model to learn these score functions from multimodal offline data, and integrate them into joint-action critics to guide policy updates toward high-reward, in-distribution regions under a shared team reward. Our approach achieves state-of-the-art performance across multiple particle environments and Multi-agent MuJoCo benchmarks consistently. To the best of our knowledge, this is the first work to explicitly address the distributional gap between offline and online MARL, paving the way for more generalizable offline policy-based MARL methods.
翻译:离线合作式多智能体强化学习因分布偏移面临独特挑战,该问题主要源于联合动作空间的高维特性以及超出分布范围的联合动作选择。本工作指出,离线多智能体强化学习的根本挑战在于合作任务的多元均衡特性,这导致产生高度多模态的联合行为策略空间与异质性质量行为数据的结合,使得单一策略正则化难以对齐一致的协调模式,进而引发策略分布偏移问题。为解决该挑战,我们设计了一种序贯得分函数分解方法,从联合行为策略中提取每个智能体的正则化信号,在分散执行约束下实现协调模态选择。进一步,我们利用灵活的扩散生成模型从多模态离线数据中学习这些得分函数,并将其集成到联合动作评论器中,引导策略在共享团队奖励下向高回报、分布内区域更新。该方法在多个粒子环境与多智能体MuJoCo基准测试中持续达到最优性能。据我们所知,这是首项明确解决离线与在线多智能体强化学习之间分布差距的工作,为更具泛化性的离线策略型多智能体强化学习方法奠定了基础。