We study a setting in which two players play a (possibly approximate) Nash equilibrium of a bimatrix game, while a learner observes only their actions and has no knowledge of the equilibrium or the underlying game. A natural question is whether the learner can rationalize the observed behavior by inferring the players' payoff functions. Rather than producing a single payoff estimate, inverse game theory aims to identify the entire set of payoffs consistent with observed behavior, enabling downstream use in, e.g., counterfactual analysis and mechanism design across applications like auctions, pricing, and security games. We focus on the problem of estimating the set of feasible payoffs with high probability and up to precision $ε$ on the Hausdorff metric. We provide the first minimax-optimal rates for both exact and approximate equilibrium play, in zero-sum as well as general-sum games. Our results provide learning-theoretic foundations for set-valued payoff inference in multi-agent environments.
翻译:我们研究一个场景:两名玩家在双矩阵博弈中玩某个(可能近似的)纳什均衡,而学习者仅观察其行动,且对均衡或底层博弈一无所知。一个自然的问题是,学习者能否通过推断玩家的收益函数来合理解释观察到的行为。逆向博弈论的目标并非生成单一收益估计,而是识别与观察行为一致的所有收益集合,从而为反事实分析、机制设计等下游应用提供支持(例如拍卖、定价和安全博弈)。我们重点研究以高概率估计可行收益集合的问题,并要求在Hausdorff度量下达到$ε$精度。我们首次为零和博弈及一般和博弈中的精确与近似均衡行为提供了极小化最优速率。我们的结果为多智能体环境中集合值收益推断奠定了学习理论基础。