Real-world competitive games, such as chess, go, or StarCraft II, rely on Elo models to measure the strength of their players. Since these games are not fully transitive, using Elo implicitly assumes they have a strong transitive component that can correctly be identified and extracted. In this study, we investigate the challenge of identifying the strength of the transitive component in games. First, we show that Elo models can fail to extract this transitive component, even in elementary transitive games. Then, based on this observation, we propose an extension of the Elo score: we end up with a disc ranking system that assigns each player two scores, which we refer to as skill and consistency. Finally, we propose an empirical validation on payoff matrices coming from real-world games played by bots and humans.
翻译:现实世界中的竞争性博弈,如国际象棋、围棋或《星际争霸II》,依赖Elo模型来衡量玩家的实力。由于这些博弈并非完全具有传递性,使用Elo模型隐含地假定它们具有较强的可被正确识别与提取的传递性成分。在本研究中,我们探讨了识别博弈中传递性成分强度的挑战。首先,我们证明即使在基本的传递性博弈中,Elo模型也可能无法提取这一传递性成分。随后,基于这一发现,我们提出了一种对Elo评分的扩展:最终形成一种圆盘排名系统,为每位玩家分配两个分数,我们称之为技能与一致性。最后,我们利用来自机器人与人类玩家参与的现实世界博弈的收益矩阵,进行了实证验证。