A dynamic mean field theory is developed for finite state and action Bayesian reinforcement learning in the large state space limit. In an analogy with statistical physics, the Bellman equation is studied as a disordered dynamical system; the Markov decision process transition probabilities are interpreted as couplings and the value functions as deterministic spins that evolve dynamically. Thus, the mean-rewards and transition probabilities are considered to be quenched random variables. The theory reveals that, under certain assumptions, the state-action values are statistically independent across state-action pairs in the asymptotic state space limit, and provides the form of the distribution exactly. The results hold in the finite and discounted infinite horizon settings, for both value iteration and policy evaluation. The state-action value statistics can be computed from a set of mean field equations, which we call dynamic mean field programming (DMFP). For policy evaluation the equations are exact. For value iteration, approximate equations are obtained by appealing to extreme value theory or bounds. The result provides analytic insight into the statistical structure of tabular reinforcement learning, for example revealing the conditions under which reinforcement learning is equivalent to a set of independent multi-armed bandit problems.
翻译:针对大状态空间极限下的有限状态与动作贝叶斯强化学习,本文发展了一种动态平均场理论。类比于统计物理,将贝尔曼方程作为无序动力系统进行研究:马尔可夫决策过程的转移概率被解释为耦合,而价值函数被视为动态演化的确定性自旋。因此,平均奖励与转移概率被视作淬火随机变量。该理论揭示,在特定假设下,状态-动作值在渐近状态空间极限中跨状态-动作对统计独立,并精确给出了其分布形式。该结论适用于有限步长与折现无限时域场景,涵盖值迭代与策略评估。状态-动作值统计量可通过一组平均场方程计算,我们称之为动态平均场规划。对于策略评估,方程为精确形式;对于值迭代,则需借助极值理论或边界获得近似方程。该结果为表格型强化学习的统计结构提供了解析洞察,例如揭示了强化学习等价于一组独立多臂赌博机问题的条件。