Ensuring safety in dynamic multi-agent systems is challenging due to limited information about the other agents. Control Barrier Functions (CBFs) are showing promise for safety assurance but current methods make strong assumptions about other agents and often rely on manual tuning to balance safety, feasibility, and performance. In this work, we delve into the problem of adaptive safe learning for multi-agent systems with CBF. We show how emergent behavior can be profoundly influenced by the CBF configuration, highlighting the necessity for a responsive and dynamic approach to CBF design. We present ASRL, a novel adaptive safe RL framework, to fully automate the optimization of policy and CBF coefficients, to enhance safety and long-term performance through reinforcement learning. By directly interacting with the other agents, ASRL learns to cope with diverse agent behaviours and maintains the cost violations below a desired limit. We evaluate ASRL in a multi-robot system and a competitive multi-agent racing scenario, against learning-based and control-theoretic approaches. We empirically demonstrate the efficacy and flexibility of ASRL, and assess generalization and scalability to out-of-distribution scenarios. Code and supplementary material are public online.
翻译:在动态多智能体系统中确保安全颇具挑战性,原因在于对其他智能体的信息有限。控制障碍函数(CBF)在安全保证方面展现出潜力,但现有方法对其他智能体做出了较强假设,且常依赖人工调参来平衡安全性、可行性与性能。本文深入研究了基于CBF的多智能体系统自适应安全学习问题。我们揭示了CBF配置如何深刻影响涌现行为,凸显了采用响应式动态方法进行CBF设计的必要性。我们提出了ASRL——一种新颖的自适应安全强化学习框架,通过强化学习全自动化优化策略与CBF系数,以增强安全性并提升长期性能。通过直接与其他智能体交互,ASRL能够学习应对多样化的智能体行为,并将代价违规控制在期望阈值以下。我们在多机器人系统与竞争性多智能体赛车场景中,将ASRL与基于学习及控制理论的方法进行了对比评估。实证结果证明了ASRL的有效性与灵活性,并评估了其在分布外场景中的泛化能力与可扩展性。代码与补充材料已公开在线。