Actor-critic methods have achieved significant success in many challenging applications. However, its finite-time convergence is still poorly understood in the most practical single-timescale form. Existing works on analyzing single-timescale actor-critic have been limited to i.i.d. sampling or tabular setting for simplicity. We investigate the more practical online single-timescale actor-critic algorithm on continuous state space, where the critic assumes linear function approximation and updates with a single Markovian sample per actor step. Previous analysis has been unable to establish the convergence for such a challenging scenario. We demonstrate that the online single-timescale actor-critic method provably finds an $\epsilon$-approximate stationary point with $\widetilde{\mathcal{O}}(\epsilon^{-2})$ sample complexity under standard assumptions, which can be further improved to $\mathcal{O}(\epsilon^{-2})$ under the i.i.d. sampling. Our novel framework systematically evaluates and controls the error propagation between the actor and critic. It offers a promising approach for analyzing other single-timescale reinforcement learning algorithms as well.
翻译:演员-评论家方法在许多具有挑战性的应用中取得了显著成功。然而,在其最实用的单时间尺度形式下,该算法的有限时间收敛性仍未被充分理解。现有关于单时间尺度演员-评论家的分析工作,为简化问题通常局限于独立同分布采样或表格设置。本文研究了连续状态空间下更实际的在线单时间尺度演员-评论家算法,其中评论家采用线性函数逼近,并在每次演员步长后基于单个马尔可夫样本进行参数更新。先前的分析未能建立这种具有挑战性场景下的收敛性。我们证明,在标准假设下,在线单时间尺度演员-评论家方法能以$\widetilde{\mathcal{O}}(\epsilon^{-2})$的样本复杂度找到一个$\epsilon$-近似平稳点;在独立同分布采样下,该复杂度可进一步改进为$\mathcal{O}(\epsilon^{-2})$。我们提出的新颖框架系统性地评估并控制了演员与评论家之间的误差传播,为分析其他单时间尺度强化学习算法提供了有前景的途径。