Online reinforcement learning in infinite-horizon Markov decision processes (MDPs) remains less theoretically and algorithmically developed than its episodic counterpart, with many algorithms suffering from high ``burn-in'' costs and failing to adapt to benign instance-specific complexity. In this work, we address these shortcomings for two infinite-horizon objectives: the classical average-reward regret and the $γ$-regret. We develop a single tractable UCB-style algorithm applicable to both settings, which achieves the first optimal variance-dependent regret guarantees. Our regret bounds in both settings take the form $\tilde{O}( \sqrt{SA\,\text{Var}} + \text{lower-order terms})$, where $S,A$ are the state and action space sizes, and $\text{Var}$ captures cumulative transition variance. This implies minimax-optimal average-reward and $γ$-regret bounds in the worst case but also adapts to easier problem instances, for example yielding nearly constant regret in deterministic MDPs. Furthermore, our algorithm enjoys significantly improved lower-order terms for the average-reward setting. With prior knowledge of the optimal bias span $\Vert h^\star\Vert_\text{sp}$, our algorithm obtains lower-order terms scaling as $\Vert h^\star\Vert_\text{sp} S^2 A$, which we prove is optimal in both $\Vert h^\star\Vert_\text{sp}$ and $A$. Without prior knowledge, we prove that no algorithm can have lower-order terms smaller than $\Vert h^\star \Vert_\text{sp}^2 S A$, and we provide a prior-free algorithm whose lower-order terms scale as $\Vert h^\star\Vert_\text{sp}^2 S^3 A$, nearly matching this lower bound. Taken together, these results completely characterize the optimal dependence on $\Vert h^\star\Vert_\text{sp}$ in both leading and lower-order terms, and reveal a fundamental gap in what is achievable with and without prior knowledge.
翻译:在线强化学习在无限期马尔可夫决策过程(MDP)中的理论与算法发展仍落后于其回合制对应方法,许多算法存在高"预热"成本且无法适应良性的实例特定复杂度。本文针对两种无限期目标——经典平均奖励遗憾与γ-遗憾——克服了上述缺陷。我们提出了一种适用于两种场景的单一易处理UCB型算法,首次实现了最优方差依赖遗憾保证。两种场景下的遗憾界均形如$\tilde{O}( \sqrt{SA\,\text{Var}} + \text{低阶项})$,其中$S,A$分别为状态与动作空间大小,$\text{Var}$刻画累积转移方差。这既在 worst-case 下达到极小极大最优的平均奖励与γ-遗憾界,又能适应更简单的问题实例,例如在确定性MDP中实现近乎恒定的遗憾。此外,我们的算法在平均奖励场景中显著改进了低阶项。若已知最优偏置跨度$\Vert h^\star\Vert_\text{sp}$,算法得到的低阶项量级为$\Vert h^\star\Vert_\text{sp} S^2 A$,我们证明这在$\Vert h^\star\Vert_\text{sp}$与$A$上均为最优。若无先验知识,我们证明任何算法的低阶项都不可能小于$\Vert h^\star \Vert_\text{sp}^2 S A$,并提供了一种无需先验知识的算法,其低阶项量级为$\Vert h^\star\Vert_\text{sp}^2 S^3 A$,几乎匹配该下界。综合来看,这些结果完整刻画了主项与低阶项对$\Vert h^\star\Vert_\text{sp}$的最优依赖关系,并揭示了有无先验知识时可达性之间的基本差距。