Constrained Markov decision processes (CMDPs) are a common way to model safety constraints in reinforcement learning. State-of-the-art methods for efficiently solving CMDPs are based on primal-dual algorithms. For these algorithms, all currently known regret bounds allow for error cancellations -- one can compensate for a constraint violation in one round with a strict constraint satisfaction in another. This makes the online learning process unsafe since it only guarantees safety for the final (mixture) policy but not during learning. As Efroni et al. (2020) pointed out, it is an open question whether primal-dual algorithms can provably achieve sublinear regret if we do not allow error cancellations. In this paper, we give the first affirmative answer. We first generalize a result on last-iterate convergence of regularized primal-dual schemes to CMDPs with multiple constraints. Building upon this insight, we propose a model-based primal-dual algorithm to learn in an unknown CMDP. We prove that our algorithm achieves sublinear regret without error cancellations.
翻译:约束马尔可夫决策过程(CMDP)是强化学习中建模安全约束的常用方法。目前高效求解CMDP的最先进方法基于原始-对偶算法。对于这些算法,所有已知的遗憾界均允许错误抵消——即某一轮的约束违反可由另一轮的严格约束满足来补偿。这使得在线学习过程不安全,因为其仅保证最终(混合)策略的安全性,而非学习过程中的安全性。正如Efroni等人(2020)指出的,若不允许错误抵消,原始-对偶算法能否可证明地实现次线性遗憾仍是一个开放问题。本文首次给出肯定答案。我们首先将正则化原始-对偶方案的最终迭代收敛结果推广至含多约束的CMDP。基于这一发现,我们提出一种基于模型的原始-对偶算法,用于在未知CMDP中学习。我们证明,该算法能在无错误抵消的情况下实现次线性遗憾。