Zero-shot coordination (ZSC) is a new challenge focusing on generalizing learned coordination skills to unseen partners. Existing methods train the ego agent with partners from pre-trained or evolving populations. The agent's ZSC capability is typically evaluated with a few evaluation partners, including human and agent, and reported by mean returns. Current evaluation methods for ZSC capability still need to improve in constructing diverse evaluation partners and comprehensively measuring the ZSC capability. We aim to create a reliable, comprehensive, and efficient evaluation method for ZSC capability. We formally define the ideal 'diversity-complete' evaluation partners and propose the best response (BR) diversity, which is the population diversity of the BRs to the partners, to approximate the ideal evaluation partners. We propose an evaluation workflow including 'diversity-complete' evaluation partners construction and a multi-dimensional metric, the Best Response Proximity (BR-Prox) metric. BR-Prox quantifies the ZSC capability as the performance similarity to each evaluation partner's approximate best response, demonstrating generalization capability and improvement potential. We re-evaluate strong ZSC methods in the Overcooked environment using the proposed evaluation workflow. Surprisingly, the results in some of the most used layouts fail to distinguish the performance of different ZSC methods. Moreover, the evaluated ZSC methods must produce more diverse and high-performing training partners. Our proposed evaluation workflow calls for a change in how we efficiently evaluate ZSC methods as a supplement to human evaluation.
翻译:零样本协调(ZSC)是一个新兴挑战,聚焦于将习得的协调技能泛化至未见过的伙伴。现有方法通过预训练或进化种群中的伙伴训练主体智能体,其ZSC能力通常借助少数评估伙伴(包括人类与智能体)进行测试,并以平均回报报告结果。当前ZSC能力评估方法在构建多样化评估伙伴及全面衡量ZSC能力方面仍有待改进。本文旨在创建一种可靠、全面且高效的ZSC能力评估方法。我们正式定义了理想的“多样性完备”评估伙伴,并提出最佳响应多样性(BR多样性)——即伙伴最佳响应种群多样性——以逼近理想评估伙伴。我们设计了一套评估工作流,涵盖“多样性完备”评估伙伴构建与多维度量指标——最佳响应邻近度(BR-Prox)指标。BR-Prox将ZSC能力量化为智能体与各评估伙伴近似最佳响应之间的性能相似度,从而展现泛化能力与改进潜力。我们利用所提评估工作流在Overcooked环境中重新评估了强ZSC方法,结果令人惊讶:在部分最常用的场景布局中,这些方法无法区分不同ZSC方法的性能差异。此外,被评估的ZSC方法未能生成足够多样且高性能的训练伙伴。我们提出的评估工作流呼吁改变当前ZSC方法的评估方式,作为人类评估的补充以提高效率。