Learning anticipation is a reasoning paradigm in multi-agent reinforcement learning, where agents, during learning, consider the anticipated learning of other agents. There has been substantial research into the role of learning anticipation in improving cooperation among self-interested agents in general-sum games. Two primary examples are Learning with Opponent-Learning Awareness (LOLA), which anticipates and shapes the opponent's learning process to ensure cooperation among self-interested agents in various games such as iterated prisoner's dilemma, and Look-Ahead (LA), which uses learning anticipation to guarantee convergence in games with cyclic behaviors. So far, the effectiveness of applying learning anticipation to fully-cooperative games has not been explored. In this study, we aim to research the influence of learning anticipation on coordination among common-interested agents. We first illustrate that both LOLA and LA, when applied to fully-cooperative games, degrade coordination among agents, causing worst-case outcomes. Subsequently, to overcome this miscoordination behavior, we propose Hierarchical Learning Anticipation (HLA), where agents anticipate the learning of other agents in a hierarchical fashion. Specifically, HLA assigns agents to several hierarchy levels to properly regulate their reasonings. Our theoretical and empirical findings confirm that HLA can significantly improve coordination among common-interested agents in fully-cooperative normal-form games. With HLA, to the best of our knowledge, we are the first to unlock the benefits of learning anticipation for fully-cooperative games.
翻译:学习预期是多智能体强化学习中的一种推理范式,其中智能体在学习过程中会考虑其他智能体的预期学习。已有大量研究探讨了学习预期在一般和博弈中提升自利智能体间合作的作用。两个典型案例包括:具有对手学习意识的智能体(LOLA),通过预期并塑造对手的学习过程,确保自利智能体在迭代囚徒困境等博弈中实现合作;以及前瞻性学习(LA),通过利用学习预期保证在具有循环行为的博弈中收敛。目前,尚未有研究探索学习预期在全合作博弈中的有效性。本研究旨在探讨学习预期对共同利益智能体间协调的影响。我们首先证明,LOLA和LA应用于全合作博弈时,会降低智能体间的协调能力,导致最坏情况结果。随后,为克服这种不协调行为,我们提出分层学习预期(HLA),即智能体以分层方式预期其他智能体的学习。具体而言,HLA将智能体分配到多个层级,以适当规范其推理过程。理论和实证结果均证实,HLA能显著提升全合作规范式博弈中共同利益智能体间的协调能力。据我们所知,通过HLA,我们首次实现了学习预期在全合作博弈中的优势。