The study of repeated interactions between a learner and a utility-maximizing optimizer has yielded deep insights into the manipulability of learning algorithms. However, existing literature primarily focuses on independent, unlinked rounds, largely ignoring the ubiquitous practical reality of budget constraints. In this paper, we study this interaction in repeated second-price auctions in a Bayesian setting between a learning agent and a strategic agent, both subject to strict budget constraints, showing that such cross-round constraints fundamentally alter the strategic landscape. First, we generalize the classic Stackelberg equilibrium to the Budgeted Stackelberg Equilibrium. We prove that an optimizer's optimal strategy in a budgeted setting requires time-multiplexing; for a $k$-dimensional budget constraint, the optimal strategy strictly decomposes into up to $k+1$ distinct phases, with each phase employing a possibly unique mixed strategy (the case of $k=0$ recovers the classic Stackelberg equilibrium where the optimizer repeatedly uses a single mixed strategy). Second, we address the intriguing question of non-manipulability. We prove that when the learner employs a standard Proportional controller (the "P" of the PID-controller) to pace their bids, the optimizer's utility is upper bounded by their objective value in the Budgeted Stackelberg Equilibrium baseline. By bounding the dynamics of the PID controller via a novel analysis, our results establish that this widely used control-theoretic heuristic is actually strategically robust.
翻译:本文研究了学习者和效用最大化优化者之间重复交互的机制,深入揭示了学习算法的可操纵性。然而,现有文献主要关注独立、无关联的轮次,很大程度上忽略了普遍存在的预算约束这一实际场景。在本文中,我们研究了贝叶斯环境下学习主体和策略主体在重复第二价格拍卖中的交互,双方都受严格预算约束,并证明了这种跨轮次约束从根本上改变了策略格局。首先,我们将经典斯塔克尔伯格均衡推广为预算斯塔克尔伯格均衡。我们证明,在预算约束下,优化者的最优策略需要采用时域复用:对于$k$维预算约束,最优策略严格分解为多达$k+1$个不同阶段,每个阶段可能采用独特的混合策略($k=0$的情况恢复为经典斯塔克尔伯格均衡,即优化者重复使用单一混合策略)。其次,我们探讨了关于不可操纵性的有趣问题。我们证明,当学习者采用标准比例控制器(PID控制器中的"P")来调整出价节奏时,优化者的收益被其上界为预算斯塔克尔伯格均衡基线中的目标值。通过新颖的分析方法界定PID控制器的动态特性,我们的结果表明,这种广泛使用的控制理论启发式方法实际上具有策略鲁棒性。