We present new learning dynamics combining (independent) log-linear learning and value iteration for stochastic games within the auxiliary stage game framework. The dynamics presented provably attain the efficient equilibrium (also known as optimal equilibrium) in identical-interest stochastic games, beyond the recent concentration of progress on provable convergence to some (possibly inefficient) equilibrium. The dynamics are also independent in the sense that agents take actions consistent with their local viewpoint to a reasonable extent rather than seeking equilibrium. These aspects can be of practical interest in the control applications of intelligent and autonomous systems. The key challenges are the convergence to an inefficient equilibrium and the non-stationarity of the environment from a single agent's viewpoint due to the adaptation of others. The log-linear update plays an important role in addressing the former. We address the latter through the play-in-episodes scheme in which the agents update their Q-function estimates only at the end of the episodes.
翻译:摘要:我们提出了一种结合(独立)对数线性学习与值迭代的新型学习动力学方法,该方法基于辅助阶段博弈框架。所提出的动力学方法在相同利益随机博弈中可证明达到高效均衡(亦称最优均衡),超越了近期仅关注可证明收敛至(可能低效)均衡的研究进展。该动力学方法具有独立性,即智能体在合理范围内基于局部视角采取行动,而非追求均衡。这些特性在智能自主系统的控制应用中具有实际意义。核心挑战在于收敛至低效均衡的可能性,以及因其他智能体适应性调整导致的单智能体视角下环境非平稳性。对数线性更新在应对前一挑战中发挥关键作用。我们通过情景式博弈机制解决后一挑战,即智能体仅在情景结束时更新其Q函数估计值。