Cooperative multi-agent reinforcement learning (MARL) for navigation enables agents to cooperate to achieve their navigation goals. Using emergent communication, agents learn a communication protocol to coordinate and share information that is needed to achieve their navigation tasks. In emergent communication, symbols with no pre-specified usage rules are exchanged, in which the meaning and syntax emerge through training. Learning a navigation policy along with a communication protocol in a MARL environment is highly complex due to the huge state space to be explored. To cope with this complexity, this work proposes a novel neural network architecture, for jointly learning an adaptive state space abstraction and a communication protocol among agents participating in navigation tasks. The goal is to come up with an adaptive abstractor that significantly reduces the size of the state space to be explored, without degradation in the policy performance. Simulation results show that the proposed method reaches a better policy, in terms of achievable rewards, resulting in fewer training iterations compared to the case where raw states or fixed state abstraction are used. Moreover, it is shown that a communication protocol emerges during training which enables the agents to learn better policies within fewer training iterations.
翻译:面向导航的多智能体协作强化学习使多个智能体能够协同完成导航目标。通过涌现式通信,智能体学习一种通信协议来协调和共享完成导航任务所需的信息。在涌现式通信中,智能体之间交换无预设使用规则的符号,其语义和句法通过训练过程自发形成。由于需要探索的状态空间极为庞大,在多智能体强化学习环境中联合学习导航策略与通信协议具有极高的复杂性。为应对这一挑战,本文提出一种新型神经网络架构,用于联合学习参与导航任务的多智能体之间的自适应状态空间抽象与通信协议。其目标是设计一种自适应抽象器,在不降低策略性能的前提下显著缩减待探索状态空间的规模。仿真结果表明,与采用原始状态或固定状态抽象方法相比,本方法在可达奖励指标上获得更优策略,且所需训练迭代次数更少。此外,实验证实训练过程中会涌现出通信协议,使智能体能够在更少的训练周期内学习到更优策略。