We revisit the standard formulation of tabular actor-critic algorithm as a two time-scale stochastic approximation with value function computed on a faster time-scale and policy computed on a slower time-scale. This emulates policy iteration. We observe that reversal of the time scales will in fact emulate value iteration and is a legitimate algorithm. We provide a proof of convergence and compare the two empirically with and without function approximation (with both linear and nonlinear function approximators) and observe that our proposed critic-actor algorithm performs on par with actor-critic in terms of both accuracy and computational effort.
翻译:我们重新审视了表格型演员-评论家算法的标准形式,将其视为一种双时间尺度随机逼近方法:价值函数在较快时间尺度上计算,而策略在较慢时间尺度上更新。这模拟了策略迭代。我们发现,颠倒这两个时间尺度实际上会模拟值迭代,并且是一种合理的算法。我们提供了收敛性证明,并在有无函数逼近(包括线性和非线性函数逼近器)的情况下对两者进行了实证比较,观察到我们提出的评论家-演员算法在准确性和计算效率方面均与演员-评论家算法表现相当。