Offline reinforcement learning allows control policies to be learned directly from data without online interaction, making it suitable for safety-critical tasks. Recent studies have applied diffusion models to offline reinforcement learning to leverage their strong capacity for modeling complex data distributions. However, existing approaches primarily focus on single-agent settings, leaving the safety challenges in multi-agent environments largely unexplored. In this work, we propose a safe offline multi-agent reinforcement learning algorithm that embeds neural individual control barrier functions into the diffusion model to enhance safety during trajectory generation, with control policies recovered through inverse dynamics. We evaluate our algorithm across diverse benchmarks, demonstrating substantial safety improvements while maintaining competitive rewards.
翻译:离线强化学习允许直接从数据中学习控制策略而无需在线交互,因此适用于安全关键型任务。近年来的研究将扩散模型应用于离线强化学习,以利用其建模复杂数据分布的强大能力。然而,现有方法主要聚焦于单智能体场景,多智能体环境中的安全挑战尚待充分探索。本文提出一种安全离线多智能体强化学习算法,该算法将神经个体控制屏障函数嵌入扩散模型,以在轨迹生成过程中增强安全性,并通过逆动力学恢复控制策略。我们在多种基准测试上评估该算法,结果表明其在保持竞争性奖励的同时大幅提升了安全性。