In many real-world settings agents engage in strategic interactions with multiple opposing agents who can employ a wide variety of strategies. The standard approach for designing agents for such settings is to compute or approximate a relevant game-theoretic solution concept such as Nash equilibrium and then follow the prescribed strategy. However, such a strategy ignores any observations of opponents' play, which may indicate shortcomings that can be exploited. We present an approach for opponent modeling in multiplayer imperfect-information games where we collect observations of opponents' play through repeated interactions. We run experiments against a wide variety of real opponents and exact Nash equilibrium strategies in three-player Kuhn poker and show that our algorithm significantly outperforms all of the agents, including the exact Nash equilibrium strategies.
翻译:在许多现实场景中,智能体需要与多个可采用多种策略的对手进行战略互动。针对此类场景设计智能体的标准方法是计算或近似相关博弈论解概念(如纳什均衡),随后遵循该策略行动。然而,这种策略会忽略对手行为的任何观察结果——这些观察可能揭示可被利用的缺陷。我们提出了一种多智能体不完全信息游戏中对手建模的方法,通过重复交互收集对手的玩法观察。我们在三人库恩扑克中针对多种真实对手和精确纳什均衡策略进行了实验,结果表明我们的算法在性能上显著超越所有智能体,包括精确纳什均衡策略。