Constrained multiagent reinforcement learning (C-MARL) is gaining importance as MARL algorithms find new applications in real-world systems ranging from energy systems to drone swarms. Most C-MARL algorithms use a primal-dual approach to enforce constraints through a penalty function added to the reward. In this paper, we study the structural effects of this penalty term on the MARL problem. First, we show that the standard practice of using the constraint function as the penalty leads to a weak notion of safety. However, by making simple modifications to the penalty term, we can enforce meaningful probabilistic (chance and conditional value at risk) constraints. Second, we quantify the effect of the penalty term on the value function, uncovering an improved value estimation procedure. We use these insights to propose a constrained multiagent advantage actor critic (C-MAA2C) algorithm. Simulations in a simple constrained multiagent environment affirm that our reinterpretation of the primal-dual method in terms of probabilistic constraints is effective, and that our proposed value estimate accelerates convergence to a safe joint policy.
翻译:约束多智能体强化学习(C-MARL)随着MARL算法在从能源系统到无人机集群等实际系统中的新应用而日益重要。大多数C-MARL算法采用原始-对偶方法,通过向奖励函数中添加惩罚项来强制执行约束。本文研究了该惩罚项对MARL问题的结构性影响。首先,我们证明将约束函数直接用作惩罚项的标准做法会导致较弱的安全概念。然而,通过对惩罚项进行简单修改,我们可以强制执行有意义的概率约束(机会约束和条件风险价值)。其次,我们量化了惩罚项对价值函数的影响,揭示了一种改进的价值估计方法。基于这些见解,我们提出了一种约束多智能体优势演员-评论家(C-MAA2C)算法。在简单约束多智能体环境中的仿真验证了我们对原始-对偶方法在概率约束框架下的重新解释是有效的,并且我们提出的价值估计能加速收敛至安全的联合策略。