Offline cooperative multi-agent reinforcement learning (MARL) faces unique challenges due to distributional shifts, particularly stemming from the high dimensionality of joint action spaces and the presence of out-of-distribution joint action selections. In this work, we highlight that a fundamental challenge in offline MARL arises from the multi-equilibrium nature of cooperative tasks, which induces a highly multimodal joint behavior policy space coupled with heterogeneous-quality behavior data. This makes it difficult for individual policy regularization to align with a consistent coordination pattern, leading to the policy distribution shift problems. To tackle this challenge, we design a sequential score function decomposition method that distills per-agent regularization signals from the joint behavior policy, which induces coordinated modality selection under decentralized execution constraints. Then we leverage a flexible diffusion-based generative model to learn these score functions from multimodal offline data, and integrate them into joint-action critics to guide policy updates toward high-reward, in-distribution regions under a shared team reward. Our approach achieves state-of-the-art performance across multiple particle environments and Multi-agent MuJoCo benchmarks consistently. To the best of our knowledge, this is the first work to explicitly address the distributional gap between offline and online MARL, paving the way for more generalizable offline policy-based MARL methods.


翻译:离线合作式多智能体强化学习因分布偏移面临独特挑战,该问题主要源于联合动作空间的高维特性以及超出分布范围的联合动作选择。本工作指出,离线多智能体强化学习的根本挑战在于合作任务的多元均衡特性,这导致产生高度多模态的联合行为策略空间与异质性质量行为数据的结合,使得单一策略正则化难以对齐一致的协调模式,进而引发策略分布偏移问题。为解决该挑战,我们设计了一种序贯得分函数分解方法,从联合行为策略中提取每个智能体的正则化信号,在分散执行约束下实现协调模态选择。进一步,我们利用灵活的扩散生成模型从多模态离线数据中学习这些得分函数,并将其集成到联合动作评论器中,引导策略在共享团队奖励下向高回报、分布内区域更新。该方法在多个粒子环境与多智能体MuJoCo基准测试中持续达到最优性能。据我们所知,这是首项明确解决离线与在线多智能体强化学习之间分布差距的工作,为更具泛化性的离线策略型多智能体强化学习方法奠定了基础。

0
下载
关闭预览

相关内容

《基于分层多智能体强化学习的逼真空战协同策略》
专知会员服务
48+阅读 · 2025年10月30日
基于多智能体强化学习的博弈综述
专知会员服务
53+阅读 · 2024年11月23日
多智能体博弈中的分布式学习: 原理与算法
专知会员服务
54+阅读 · 2024年6月13日
基于学习机制的多智能体强化学习综述
专知会员服务
64+阅读 · 2024年4月16日
基于多智能体强化学习的协同目标分配
专知会员服务
142+阅读 · 2023年9月5日
基于多智能体深度强化学习的体系任务分配方法
专知会员服务
159+阅读 · 2023年5月4日
专知会员服务
172+阅读 · 2021年8月3日
「基于通信的多智能体强化学习」 进展综述
【综述】多智能体强化学习算法理论研究
深度强化学习实验室
16+阅读 · 2020年9月9日
多智能体强化学习(MARL)近年研究概览
PaperWeekly
38+阅读 · 2020年3月15日
DeepMind:用PopArt进行多任务深度强化学习
论智
30+阅读 · 2018年9月14日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
24+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
VIP会员
最新内容
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 今天4:08
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
“史诗怒火”行动:现代多域作战的重要节点
专知会员服务
8+阅读 · 7月30日
《下一代无线网络中的多无人机通信资源管理》
相关VIP内容
《基于分层多智能体强化学习的逼真空战协同策略》
专知会员服务
48+阅读 · 2025年10月30日
基于多智能体强化学习的博弈综述
专知会员服务
53+阅读 · 2024年11月23日
多智能体博弈中的分布式学习: 原理与算法
专知会员服务
54+阅读 · 2024年6月13日
基于学习机制的多智能体强化学习综述
专知会员服务
64+阅读 · 2024年4月16日
基于多智能体强化学习的协同目标分配
专知会员服务
142+阅读 · 2023年9月5日
基于多智能体深度强化学习的体系任务分配方法
专知会员服务
159+阅读 · 2023年5月4日
专知会员服务
172+阅读 · 2021年8月3日
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
24+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员