Cooperative multi-agent reinforcement learning (MARL) for navigation enables agents to cooperate to achieve their navigation goals. Using emergent communication, agents learn a communication protocol to coordinate and share information that is needed to achieve their navigation tasks. In emergent communication, symbols with no pre-specified usage rules are exchanged, in which the meaning and syntax emerge through training. Learning a navigation policy along with a communication protocol in a MARL environment is highly complex due to the huge state space to be explored. To cope with this complexity, this work proposes a novel neural network architecture, for jointly learning an adaptive state space abstraction and a communication protocol among agents participating in navigation tasks. The goal is to come up with an adaptive abstractor that significantly reduces the size of the state space to be explored, without degradation in the policy performance. Simulation results show that the proposed method reaches a better policy, in terms of achievable rewards, resulting in fewer training iterations compared to the case where raw states or fixed state abstraction are used. Moreover, it is shown that a communication protocol emerges during training which enables the agents to learn better policies within fewer training iterations.
翻译:合作式多智能体强化学习(MARL)在导航任务中使智能体能够相互协作以实现其导航目标。通过涌现通信,智能体学习一种通信协议来协调并共享完成导航任务所需的信息。在涌现通信中,智能体之间交换没有预先指定使用规则的符号,其含义和句法通过训练过程自然形成。由于待探索的状态空间极其庞大,在MARL环境中同时学习导航策略与通信协议具有高度复杂性。为应对这一挑战,本文提出了一种新型神经网络架构,用于联合学习自适应状态空间抽象及参与导航任务的智能体间的通信协议。该架构旨在设计一种自适应抽象器,在不降低策略性能的前提下显著缩减待探索状态空间的规模。仿真结果表明,相较于使用原始状态或固定状态抽象的方法,所提方法在可获取奖励方面能达到更优策略,同时减少训练迭代次数。此外,研究显示训练过程中涌现出的通信协议能使智能体在更少训练轮次内学习到更优策略。