How can the system operator learn an incentive mechanism that achieves social optimality based on limited information about the agents' behavior, who are dynamically updating their strategies? To answer this question, we propose an \emph{adaptive} incentive mechanism. This mechanism updates the incentives of agents based on the feedback of each agent's externality, evaluated as the difference between the player's marginal cost and society's marginal cost at each time step. The proposed mechanism updates the incentives on a slower timescale compared to the agents' learning dynamics, resulting in a two-timescale coupled dynamical system. Notably, this mechanism is agnostic to the specific learning dynamics used by agents to update their strategies. We show that any fixed point of this adaptive incentive mechanism corresponds to the optimal incentive mechanism, ensuring that the Nash equilibrium coincides with the socially optimal strategy. Additionally, we provide sufficient conditions that guarantee the convergence of the adaptive incentive mechanism to a fixed point. Our results apply to both atomic and non-atomic games. To demonstrate the effectiveness of our proposed mechanism, we verify the convergence conditions in two practically relevant games: atomic networked quadratic aggregative games and non-atomic network routing games.
翻译:系统操作者如何基于对智能体行为的有限信息,设计出能够实现社会最优的激励机制?这些智能体正在动态更新其策略。为回答这一问题,我们提出一种自适应激励机制。该机制根据每个智能体外部性的反馈来更新其激励值,其中外部性通过每个时间步中参与者的边际成本与社会边际成本之间的差值来评估。所提出的机制在比智能体学习动态更慢的时间尺度上更新激励,从而形成一个双时间尺度耦合动力系统。值得注意的是,该机制不依赖于智能体更新策略时所采用的具体学习动态。我们证明该自适应激励机制的任意不动点均对应于最优激励机制,确保纳什均衡与社会最优策略相一致。此外,我们提供了保证自适应激励机制收敛至不动点的充分条件。我们的结果同时适用于原子博弈与非原子博弈。为验证所提出机制的有效性,我们在两个具有实际意义的博弈中验证了收敛条件:原子网络二次聚合博弈与非原子网络路由博弈。