Many real-world games suffer from information asymmetry: one player is only aware of their own payoffs while the other player has the full game information. Examples include the critical domain of security games and adversarial multi-agent reinforcement learning. Information asymmetry renders traditional solution concepts such as Strong Stackelberg Equilibrium (SSE) and Robust-Optimization Equilibrium (ROE) inoperative. We propose a novel solution concept called VISER (Victim Is Secure, Exploiter best-Responds). VISER enables an external observer to predict the outcome of such games. In particular, for security applications, VISER allows the victim to better defend itself while characterizing the most damaging attacks available to the attacker. We show that each player's VISER strategy can be computed independently in polynomial time using linear programming (LP). We also extend VISER to its Markov-perfect counterpart for Markov games, which can be solved efficiently using a series of LPs.
翻译:许多现实博弈面临信息不对称问题:一方仅知晓自身收益,而另一方掌握完全的博弈信息。典型例子包括关键领域的安全博弈和对抗性多智能体强化学习。信息不对称导致强斯塔克尔伯格均衡(SSE)和鲁棒优化均衡(ROE)等传统解概念失效。我们提出了一种新型解概念——VISER(受害者安全、利用者最优响应)。VISER使外部观察者能够预测此类博弈的结果。特别地,在安全应用中,VISER让受害者更好地进行防御,同时刻画攻击方可实施的最具破坏性攻击。我们证明每位博弈参与方的VISER策略可通过线性规划(LP)在多项式时间内独立计算。此外,我们将VISER扩展至马尔可夫博弈的马尔可夫完美对应形式,该扩展可通过一系列线性规划高效求解。