While Nash equilibrium has emerged as the central game-theoretic solution concept, many important games contain several Nash equilibria and we must determine how to select between them in order to create real strategic agents. Several Nash equilibrium refinement concepts have been proposed and studied for sequential imperfect-information games, the most prominent being trembling-hand perfect equilibrium, quasi-perfect equilibrium, and recently one-sided quasi-perfect equilibrium. These concepts are robust to certain arbitrarily small mistakes, and are guaranteed to always exist; however, we argue that neither of these is the correct concept for developing strong agents in sequential games of imperfect information. We define a new equilibrium refinement concept for extensive-form games called observable perfect equilibrium in which the solution is robust over trembles in publicly-observable action probabilities (not necessarily over all action probabilities that may not be observable by opposing players). Observable perfect equilibrium correctly captures the assumption that the opponent is playing as rationally as possible given mistakes that have been observed (while previous solution concepts do not). We prove that observable perfect equilibrium is always guaranteed to exist, and demonstrate that it leads to a different solution than the prior extensive-form refinements in no-limit poker. We expect observable perfect equilibrium to be a useful equilibrium refinement concept for modeling many important imperfect-information games of interest in artificial intelligence.
翻译:尽管纳什均衡已成为博弈论的核心解概念,但许多重要博弈中存在多个纳什均衡,为构建真实策略智能体,我们必须确定如何在其中进行选择。针对序贯不完美信息博弈,学界已提出并研究多种纳什均衡精炼概念,其中最著名的是颤抖手完美均衡、拟完美均衡以及近期提出的单边拟完美均衡。这些概念对特定任意小错误具有鲁棒性且保证始终存在;然而我们认为,对于在不完美信息序贯博弈中开发强性能智能体而言,这些概念均非正确选择。我们为扩展式博弈定义了一种新的均衡精炼概念——可观测完美均衡,其解对公开可观测行动概率的颤抖具有鲁棒性(不要求对所有可能不被对手观测到的行动概率都鲁棒)。可观测完美均衡正确刻画了如下假设:对手在观测到已发生的错误后,会尽可能理性地行动(而之前解概念未体现此特性)。我们证明可观测完美均衡始终保证存在,并论证其在无上限扑克中会导出不同于先前扩展式精炼概念的解。我们预期可观测完美均衡将成为人工智能领域建模诸多重要不完美信息博弈的有用均衡精炼概念。