Online reinforcement learning in infinite-horizon Markov decision processes (MDPs) remains less theoretically and algorithmically developed than its episodic counterpart, with many algorithms suffering from high ``burn-in'' costs and failing to adapt to benign instance-specific complexity. In this work, we address these shortcomings for two infinite-horizon objectives: the classical average-reward regret and the $γ$-regret. We develop a single tractable UCB-style algorithm applicable to both settings, which achieves the first optimal variance-dependent regret guarantees. Our regret bounds in both settings take the form $\tilde{O}( \sqrt{SA\,\text{Var}} + \text{lower-order terms})$, where $S,A$ are the state and action space sizes, and $\text{Var}$ captures cumulative transition variance. This implies minimax-optimal average-reward and $γ$-regret bounds in the worst case but also adapts to easier problem instances, for example yielding nearly constant regret in deterministic MDPs. Furthermore, our algorithm enjoys significantly improved lower-order terms for the average-reward setting. With prior knowledge of the optimal bias span $\Vert h^\star\Vert_\text{sp}$, our algorithm obtains lower-order terms scaling as $\Vert h^\star\Vert_\text{sp} S^2 A$, which we prove is optimal in both $\Vert h^\star\Vert_\text{sp}$ and $A$. Without prior knowledge, we prove that no algorithm can have lower-order terms smaller than $\Vert h^\star \Vert_\text{sp}^2 S A$, and we provide a prior-free algorithm whose lower-order terms scale as $\Vert h^\star\Vert_\text{sp}^2 S^3 A$, nearly matching this lower bound. Taken together, these results completely characterize the optimal dependence on $\Vert h^\star\Vert_\text{sp}$ in both leading and lower-order terms, and reveal a fundamental gap in what is achievable with and without prior knowledge.


翻译:在线强化学习在无限期马尔可夫决策过程(MDP)中的理论与算法发展仍落后于其回合制对应方法,许多算法存在高"预热"成本且无法适应良性的实例特定复杂度。本文针对两种无限期目标——经典平均奖励遗憾与γ-遗憾——克服了上述缺陷。我们提出了一种适用于两种场景的单一易处理UCB型算法,首次实现了最优方差依赖遗憾保证。两种场景下的遗憾界均形如$\tilde{O}( \sqrt{SA\,\text{Var}} + \text{低阶项})$,其中$S,A$分别为状态与动作空间大小,$\text{Var}$刻画累积转移方差。这既在 worst-case 下达到极小极大最优的平均奖励与γ-遗憾界,又能适应更简单的问题实例,例如在确定性MDP中实现近乎恒定的遗憾。此外,我们的算法在平均奖励场景中显著改进了低阶项。若已知最优偏置跨度$\Vert h^\star\Vert_\text{sp}$,算法得到的低阶项量级为$\Vert h^\star\Vert_\text{sp} S^2 A$,我们证明这在$\Vert h^\star\Vert_\text{sp}$与$A$上均为最优。若无先验知识,我们证明任何算法的低阶项都不可能小于$\Vert h^\star \Vert_\text{sp}^2 S A$,并提供了一种无需先验知识的算法,其低阶项量级为$\Vert h^\star\Vert_\text{sp}^2 S^3 A$,几乎匹配该下界。综合来看,这些结果完整刻画了主项与低阶项对$\Vert h^\star\Vert_\text{sp}$的最优依赖关系,并揭示了有无先验知识时可达性之间的基本差距。

0
下载
关闭预览

相关内容

【NeurIPS2023】强化学习中的概率推理:正确的方法
专知会员服务
28+阅读 · 2023年11月25日
基于模型的强化学习综述
专知会员服务
48+阅读 · 2023年1月9日
《分布式多智能体强化学习的编码》加州大学等
专知会员服务
57+阅读 · 2022年11月2日
【普林斯顿-Mengdi Wang】强化学习统计复杂度,35页ppt
专知会员服务
21+阅读 · 2020年11月15日
548页MIT强化学习教程,收藏备用【PDF下载】
机器学习算法与Python学习
17+阅读 · 2018年10月11日
【强化学习】强化学习/增强学习/再励学习介绍
产业智能官
10+阅读 · 2018年2月23日
【论文】变分推断(Variational inference)的总结
机器学习研究会
39+阅读 · 2017年11月16日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Arxiv
0+阅读 · 4月27日
VIP会员
相关主题
最新内容
《边缘计算关键技术分析及美军作战实践应用》
专知会员服务
0+阅读 · 今天14:08
边缘计算的军事应用
专知会员服务
1+阅读 · 今天13:50
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
4+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
10+阅读 · 8月7日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员