In this work, we propose a framework that combines multi-agent reinforcement learning (MARL) with model-based control to achieve safe, dynamically feasible actions in cooperative multi-agent tasks. Multi-agent reinforcement learning provides the advantage of learning cooperative policies for multi-agent teams from discrete non-differentiable rewards in a long planning horizon. Model-predictive control is robust and offers safe, dynamically feasible actions in a fast replanning framework for short horizons. We propose an algorithm that extends actor-critic model predictive control for MARL which we refer to as multi-agent actor-critic model predictive control (MA-AC-MPC). We demonstrate the capabilities of this algorithm by applying it to a multi-agent pursuit-evasion scenario. Specifically, we compare the evader team's strategy using the MA-AC-MPC model and a multi-layer perceptron model (MA-AC-MLP). The pursuer team uses augmented proportional navigation as it is accepted as an advanced adversarial control law. We also provide an example with a heterogeneous environment where a drone and omni-wheeled rover cooperate to achieve repeatable and successful landing with 100% success rate in hardware for MA-AC-MPC compared to 60% for MA-AC-MLP. We demonstrate the robustness of the proposed MA-AC-MPC algorithm in hardware for both environments.
翻译:本研究提出一种融合多智能体强化学习(MARL)与基于模型控制的框架,用于在协同多智能体任务中实现安全且动态可行的动作。多智能体强化学习具有从长时域规划中基于离散不可微奖励学习多智能体团队协同策略的优势。模型预测控制则具有鲁棒性,能在短时域快速重规划框架中提供安全且动态可行的动作。我们提出一种面向MARL的基于演员-评论家的模型预测控制扩展算法,并称之为多智能体演员-评论家模型预测控制(MA-AC-MPC)。通过将该算法应用于多智能体追捕-逃脱场景,我们展示了其能力。具体而言,我们比较了逃逸方采用MA-AC-MPC模型与多层感知机模型(MA-AC-MLP)的策略。追捕方采用增强比例导引法,因其被公认为先进的对抗性控制律。我们还给出了一个异质环境示例:在该环境中,无人机与全向轮式移动机器人协同工作,MA-AC-MPC在硬件中实现了100%的可重复成功着陆率,而MA-AC-MLP的成功率为60%。我们通过两个环境下的硬件实验验证了所提MA-AC-MPC算法的鲁棒性。