Ad-hoc team cooperation is the problem of cooperating with other players that have not been seen in the learning process. Recently, this problem has been considered in the context of Hanabi, which requires cooperation without explicit communication with the other players. While in self-play strategies cooperating on reinforcement learning (RL) process has shown success, there is the problem of failing to cooperate with other unseen agents after the initial learning is completed. In this paper, we categorize the results of ad-hoc team cooperation into Failure, Success, and Synergy and analyze the associated failures. First, we confirm that agents learning via RL converge to one strategy each, but not necessarily the same strategy and that these agents can deploy different strategies even though they utilize the same hyperparameters. Second, we confirm that the larger the behavioral difference, the more pronounced the failure of ad-hoc team cooperation, as demonstrated using hierarchical clustering and Pearson correlation. We confirm that such agents are grouped into distinctly different groups through hierarchical clustering, such that the correlation between behavioral differences and ad-hoc team performance is -0.978. Our results improve understanding of key factors to form successful ad-hoc team cooperation in multi-player games.
翻译:特设团队合作是指与在学习过程中未遇到的其他玩家进行协作的问题。近年来,这一问题在Hanabi游戏背景下得到关注,该游戏要求在不与其他玩家进行显式沟通的情况下实现合作。尽管基于强化学习(RL)的自对弈策略在协作中取得了成功,但在初始学习完成后,与未见过的其他智能体合作时仍存在失败问题。本文将对特设团队合作的结果分为失败、成功和协同三类,并分析相关失败原因。首先,我们证实通过RL学习的智能体会各自收敛到一种策略,但未必是相同策略,且即使使用相同超参数,这些智能体也可能采用不同策略。其次,通过层次聚类和皮尔逊相关性分析,我们证实行为差异越大,特设团队合作的失败越显著。研究显示,这类智能体通过层次聚类被划分为明显不同的组别,且行为差异与特设团队绩效之间的相关系数为-0.978。这些结果加深了我们对多人游戏中成功实现特设团队合作关键因素的理解。