Finding a balance between collaboration and competition is crucial for artificial agents in many real-world applications. We investigate this using a Multi-Agent Reinforcement Learning (MARL) setup on the back of a high-impact problem. The accumulation and yearly growth of plastic in the ocean cause irreparable damage to many aspects of oceanic health and the marina system. To prevent further damage, we need to find ways to reduce macroplastics from known plastic patches in the ocean. Here we propose a Graph Neural Network (GNN) based communication mechanism that increases the agents' observation space. In our custom environment, agents control a plastic collecting vessel. The communication mechanism enables agents to develop a communication protocol using a binary signal. While the goal of the agent collective is to clean up as much as possible, agents are rewarded for the individual amount of macroplastics collected. Hence agents have to learn to communicate effectively while maintaining high individual performance. We compare our proposed communication mechanism with a multi-agent baseline without the ability to communicate. Results show communication enables collaboration and increases collective performance significantly. This means agents have learned the importance of communication and found a balance between collaboration and competition.
翻译:在众多真实世界应用中,在协作与竞争之间找到平衡对人工智能体而言至关重要。我们基于一个高影响力问题,采用多智能体强化学习(MARL)框架对此展开研究。海洋中塑料的累积与逐年增长对海洋健康及海洋生态系统的诸多方面造成不可逆损害。为阻止进一步破坏,我们需要找到从已知海洋塑料斑块中减少宏观塑料的方法。本文提出一种基于图神经网络(GNN)的通信机制,以扩展智能体的观测空间。在我们的定制环境中,智能体操控塑料收集船。该通信机制使智能体能够通过二进制信号发展出通信协议。尽管智能体集体的目标是尽可能多地清理塑料,但每个智能体因其个体收集的宏观塑料量而获得奖励。因此,智能体需要在保持高个体性能的同时,学会有效沟通。我们将所提出的通信机制与不具备通信能力的多智能体基线进行对比。结果表明,通信能够促进协作并显著提升集体性能。这意味着智能体已学会沟通的重要性,并在协作与竞争之间找到了平衡。