It is shown in recent studies that in a Stackelberg game the follower can manipulate the leader by deviating from their true best-response behavior. Such manipulations are computationally tractable and can be highly beneficial for the follower. Meanwhile, they may result in significant payoff losses for the leader, sometimes completely defeating their first-mover advantage. A warning to commitment optimizers, the risk these findings indicate appears to be alleviated to some extent by a strict information advantage the manipulations rely on. That is, the follower knows the full information about both players' payoffs whereas the leader only knows their own payoffs. In this paper, we study the manipulation problem with this information advantage relaxed. We consider the scenario where the follower is not given any information about the leader's payoffs to begin with but has to learn to manipulate by interacting with the leader. The follower can gather necessary information by querying the leader's optimal commitments against contrived best-response behaviors. Our results indicate that the information advantage is not entirely indispensable to the follower's manipulations: the follower can learn the optimal way to manipulate in polynomial time with polynomially many queries of the leader's optimal commitment.
翻译:近年来研究表明,在斯塔克尔伯格博弈中,追随者可以通过偏离其真实最佳反应行为来操纵领导者。这种操纵在计算上易于实现,且对追随者可能极为有利。同时,它们可能导致领导者遭受显著收益损失,有时甚至完全抵消其先发优势。作为对承诺优化器的警告,这些发现所揭示的风险似乎在一定程度上被操纵所依赖的严格信息优势所缓解。即,追随者掌握双方收益的完整信息,而领导者仅了解自身收益。本文研究了放松这种信息优势的操纵问题。我们考虑以下场景:追随者最初不掌握领导者收益的任何信息,但需通过与领导者交互来学习操纵策略。追随者可通过查询领导者针对人为最佳反应行为的最优承诺来收集必要信息。我们的结果表明,信息优势对追随者的操纵并非完全不可或缺:追随者可通过多次查询领导者最优承诺,在多项式时间内以多项式复杂度习得最优操纵方式。