In this paper, we establish last-iterate convergence rates for off-policy actor--critic methods in reinforcement learning. In particular, under a single-loop, single-timescale implementation and a broad class of policy updates, including approximate policy iteration and natural policy gradient methods, we prove the first $\tilde{\mathcal{O}}(ε^{-2})$ sample complexity guarantee for finding an $ε$-optimal policy under minimal assumptions, namely, the existence of a policy that induces an irreducible Markov chain. This stands in stark contrast to the existing literature, where an $\tilde{\mathcal{O}}(ε^{-2})$ sample complexity is achieved only through nested-loop updates and/or under strong, algorithm-dependent assumptions on the policies, such as uniform mixing and uniform exploration. Technically, to address the challenges posed by the coupled update equations arising from the single-loop implementation, as well as the potentially unbounded iterates induced by off-policy learning, our analysis is based on a coupled Lyapunov drift framework. Specifically, we establish a geometric convergence rate for the actor and an $\tilde{\mathcal{O}}(1/T)$ convergence rate for the critic, and combine the two Lyapunov drift inequalities through a cross-domination property. We believe this analytical framework is of independent interest and may be applicable to other coupled iterative algorithms with unbounded


翻译:本文建立了强化学习中离策略演员-评论家方法的末次迭代收敛速率。具体而言,在单循环、单时间尺度实现及包括近似策略迭代和自然策略梯度方法在内的广泛策略更新类别下,我们证明了在最小假设——即存在能生成不可约马尔可夫链的策略——下寻找 $ε$-最优策略的首个 $\tilde{\mathcal{O}}(ε^{-2})$ 样本复杂度保证。这与现有文献形成鲜明对比,后者仅通过嵌套循环更新和/或在强算法依赖假设(如均匀混合和均匀探索)下才能实现 $\tilde{\mathcal{O}}(ε^{-2})$ 的样本复杂度。技术上,为应对单循环实现中耦合更新方程以及离策略学习可能导致的无界迭代量所带来的挑战,我们的分析基于一个耦合李雅普诺夫漂移框架。具体而言,我们建立了演员的几何收敛速率和评论家的 $\tilde{\mathcal{O}}(1/T)$ 收敛速率,并通过交叉支配性质将两个李雅普诺夫漂移不等式相结合。我们相信这一分析框架具有独立意义,并可能适用于其他具有无界迭代量的耦合迭代算法。

0
下载
关闭预览

相关内容

【NeurIPS2023】强化学习中的概率推理:正确的方法
专知会员服务
28+阅读 · 2023年11月25日
【ICML2022】基于少样本策略泛化的决策Transformer
专知会员服务
37+阅读 · 2022年7月11日
【NeurIPS 2021】设置多智能体策略梯度的方差
专知会员服务
21+阅读 · 2021年10月24日
专知会员服务
44+阅读 · 2020年9月25日
机器学习中的最优化算法总结
人工智能前沿讲习班
22+阅读 · 2019年3月22日
【论文】变分推断(Variational inference)的总结
机器学习研究会
39+阅读 · 2017年11月16日
各种相似性度量及Python实现
机器学习算法与Python学习
11+阅读 · 2017年7月6日
国家自然科学基金
44+阅读 · 2015年12月31日
国家自然科学基金
21+阅读 · 2015年12月31日
国家自然科学基金
24+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Arxiv
0+阅读 · 6月16日
VIP会员
最新内容
综述 | 面向大模型智能体的图结构个性化记忆
专知会员服务
2+阅读 · 9月10日
人工智能与未来空战管理
专知会员服务
6+阅读 · 9月9日
机器的崛起:美海军陆战队组建机器人营思考
专知会员服务
9+阅读 · 9月8日
无面之战:人工智能如何重绘权力版图
专知会员服务
6+阅读 · 9月8日
相关基金
国家自然科学基金
44+阅读 · 2015年12月31日
国家自然科学基金
21+阅读 · 2015年12月31日
国家自然科学基金
24+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员