In this paper, we establish last-iterate convergence rates for off-policy actor--critic methods in reinforcement learning. In particular, under a single-loop, single-timescale implementation and a broad class of policy updates, including approximate policy iteration and natural policy gradient methods, we prove the first $\tilde{\mathcal{O}}(ε^{-2})$ sample complexity guarantee for finding an $ε$-optimal policy under minimal assumptions, namely, the existence of a policy that induces an irreducible Markov chain. This stands in stark contrast to the existing literature, where an $\tilde{\mathcal{O}}(ε^{-2})$ sample complexity is achieved only through nested-loop updates and/or under strong, algorithm-dependent assumptions on the policies, such as uniform mixing and uniform exploration. Technically, to address the challenges posed by the coupled update equations arising from the single-loop implementation, as well as the potentially unbounded iterates induced by off-policy learning, our analysis is based on a coupled Lyapunov drift framework. Specifically, we establish a geometric convergence rate for the actor and an $\tilde{\mathcal{O}}(1/T)$ convergence rate for the critic, and combine the two Lyapunov drift inequalities through a cross-domination property. We believe this analytical framework is of independent interest and may be applicable to other coupled iterative algorithms with unbounded
翻译:本文建立了强化学习中离策略演员-评论家方法的末次迭代收敛速率。具体而言,在单循环、单时间尺度实现及包括近似策略迭代和自然策略梯度方法在内的广泛策略更新类别下,我们证明了在最小假设——即存在能生成不可约马尔可夫链的策略——下寻找 $ε$-最优策略的首个 $\tilde{\mathcal{O}}(ε^{-2})$ 样本复杂度保证。这与现有文献形成鲜明对比,后者仅通过嵌套循环更新和/或在强算法依赖假设(如均匀混合和均匀探索)下才能实现 $\tilde{\mathcal{O}}(ε^{-2})$ 的样本复杂度。技术上,为应对单循环实现中耦合更新方程以及离策略学习可能导致的无界迭代量所带来的挑战,我们的分析基于一个耦合李雅普诺夫漂移框架。具体而言,我们建立了演员的几何收敛速率和评论家的 $\tilde{\mathcal{O}}(1/T)$ 收敛速率,并通过交叉支配性质将两个李雅普诺夫漂移不等式相结合。我们相信这一分析框架具有独立意义,并可能适用于其他具有无界迭代量的耦合迭代算法。