While there is an extensive body of work characterizing the sample complexity of discounted cumulative-reward MDPs, finite sample analyses for average-reward MDPs have been limited, and most existing works rely on restrictive assumptions such as ergodicity or access to a generative model. In this work, we establish the first finite sample complexity guarantees from a single trajectory for weakly communicating average-reward MDPs. To this end, we study the dynamics of a single trajectory in weakly communicating MDPs and based on this analysis, we develop novel model-free methods. Notably, our value-based and policy-based methods provide finite sample complexity guarantees of $\widetilde{O}(1/\varepsilon^2)$ and $\widetilde{O}(1/\varepsilon^4)$ from a single trajectory in weakly communicating MDPs, respectively. Furthermore, we introduce the first model-free method that requires no prior knowledge of problem-dependent quantities for communicating MDPs.
翻译:尽管已有大量研究刻画了折扣累积奖励马尔可夫决策过程的样本复杂度,但针对平均奖励马尔可夫决策过程的有限样本分析仍较为有限,且现有工作大多依赖于诸如遍历性或生成模型访问等限制性假设。在本工作中,我们首次为弱连通平均奖励马尔可夫决策过程建立了基于单条轨迹的有限样本复杂度保证。为此,我们研究了弱连通马尔可夫决策过程中单条轨迹的动态特性,并基于该分析开发了新颖的无模型方法。值得注意的是,我们提出的基于价值的方法和基于策略的方法在弱连通马尔可夫决策过程中,分别实现了来自单条轨迹的$\widetilde{O}(1/\varepsilon^2)$和$\widetilde{O}(1/\varepsilon^4)$有限样本复杂度保证。此外,我们还首次引入了一种无需预先知晓问题相关参数的无模型方法,适用于连通马尔可夫决策过程。