Multi-Agent Reinforcement Learning (MARL) has become a promising solution for constructing a multi-agent autonomous driving system (MADS) in complex and dense scenarios. But most methods consider agents acting selfishly, which leads to conflict behaviors. Some existing works incorporate the concept of social value orientation (SVO) to promote coordination, but they lack the knowledge of other agents' SVOs, resulting in conservative maneuvers. In this paper, we aim to tackle the mentioned problem by enabling the agents to understand other agents' SVOs. To accomplish this, we propose a two-stage system framework. Firstly, we train a policy by allowing the agents to share their ground truth SVOs to establish a coordinated traffic flow. Secondly, we develop a recognition network that estimates agents' SVOs and integrates it with the policy trained in the first stage. Experiments demonstrate that our developed method significantly improves the performance of the driving policy in MADS compared to two state-of-the-art MARL algorithms.
翻译:多智能体强化学习(MARL)已成为在复杂密集场景中构建多智能体自主驾驶系统(MADS)的一种有前景的解决方案。但大多数方法假设智能体自私地行动,这会导致冲突行为。部分现有研究引入社会价值取向(SVO)概念来促进协调,但它们缺乏对其他智能体SVO的了解,从而产生保守的驾驶策略。本文旨在通过使智能体能够理解其他智能体的SVO来解决上述问题。为此,我们提出了一种两阶段系统框架。首先,我们通过允许智能体共享其真实SVO来训练一个策略,以建立协调的交通流。其次,我们开发了一个识别网络来估计智能体的SVO,并将其与第一阶段训练的策略集成。实验表明,与两种最先进的MARL算法相比,我们开发的方法显著提升了MADS中驾驶策略的性能。