By enabling agents to communicate, recent cooperative multi-agent reinforcement learning (MARL) methods have demonstrated better task performance and more coordinated behavior. Most existing approaches facilitate inter-agent communication by allowing agents to send messages to each other through free communication channels, i.e., cheap talk channels. Current methods require these channels to be constantly accessible and known to the agents a priori. In this work, we lift these requirements such that the agents must discover the cheap talk channels and learn how to use them. Hence, the problem has two main parts: cheap talk discovery (CTD) and cheap talk utilization (CTU). We introduce a novel conceptual framework for both parts and develop a new algorithm based on mutual information maximization that outperforms existing algorithms in CTD/CTU settings. We also release a novel benchmark suite to stimulate future research in CTD/CTU.
翻译:通过使智能体能够通信,近年来的合作型多智能体强化学习方法展现了更优的任务执行效果与更协调的行为。现有方法大多通过允许智能体经由免费通信渠道(即廉价交流信道)相互发送消息来促进智能体间通信。当前方法要求这些信道必须始终保持可用,且智能体需预先知晓这些信道。本研究放宽了这些限制条件,使智能体必须自行发现廉价交流信道并学习如何利用它们。因此,该问题包含两个主要部分:廉价交流发现(CTD)与廉价交流利用(CTU)。我们为这两部分提出了新的概念框架,并基于互信息最大化开发了一种新算法,该算法在CTD/CTU设置下优于现有算法。我们还发布了一套新型基准测试集,以促进CTD/CTU领域的未来研究。