This paper presents LEMURS, an algorithm for learning scalable multi-robot control policies from cooperative task demonstrations. We propose a port-Hamiltonian description of the multi-robot system to exploit universal physical constraints in interconnected systems and achieve closed-loop stability. We represent a multi-robot control policy using an architecture that combines self-attention mechanisms and neural ordinary differential equations. The former handles time-varying communication in the robot team, while the latter respects the continuous-time robot dynamics. Our representation is distributed by construction, enabling the learned control policies to be deployed in robot teams of different sizes. We demonstrate that LEMURS can learn interactions and cooperative behaviors from demonstrations of multi-agent navigation and flocking tasks.
翻译:本文提出LEMURS算法,用于从协作任务演示中学习可扩展的多机器人控制策略。我们采用端口-哈密顿描述来描述多机器人系统,以利用互联系统中通用的物理约束并实现闭环稳定性。我们使用融合自注意力机制与神经常微分方程的架构来表示多机器人控制策略:前者处理机器人团队时变通信,后者尊重连续时间机器人动力学特性。该表示天然具有分布式特性,使得学习到的控制策略能够部署到不同规模的机器人团队中。我们通过多智能体导航与集群任务的演示证明,LEMURS能从演示中学习交互行为与协作模式。