Distributional reinforcement learning (DRL), which cares about the full distribution of returns instead of just the mean, has achieved empirical success in various domains. One of the core tasks in the field of DRL is distributional policy evaluation, which involves estimating the return distribution $\eta^\pi$ for a given policy $\pi$. A distributional temporal difference (TD) algorithm has been accordingly proposed, which is an extension of the temporal difference algorithm in the classic RL literature. In the tabular case, \citet{rowland2018analysis} and \citet{rowland2023analysis} proved the asymptotic convergence of two instances of distributional TD, namely categorical temporal difference algorithm (CTD) and quantile temporal difference algorithm (QTD), respectively. In this paper, we go a step further and analyze the finite-sample performance of distributional TD. To facilitate theoretical analysis, we propose non-parametric distributional TD algorithm (NTD). For a $\gamma$-discounted infinite-horizon tabular Markov decision process with state space $S$ and action space $A$, we show that in the case of NTD we need $\wtilde O\prn{\frac{1}{\varepsilon^{2p}(1-\gamma)^{2p+2}}}$ iterations to achieve an $\varepsilon$-optimal estimator with high probability, when the estimation error is measured by the $p$-Wasserstein distance. Under some mild assumptions, $\wtilde O\prn{\frac{1}{\varepsilon^{2}(1-\gamma)^{4}}}$ iterations suffices to ensure the Kolmogorov-Smirnov distance between the NTD estimator $\hat\eta^\pi$ and $\eta^\pi$ less than $\varepsilon$ with high probability. And we revisit CTD, showing that the same non-asymptotic convergence bounds hold for CTD in the case of the $p$-Wasserstein distance.
翻译:分布强化学习(DRL)不仅关注收益的均值,更重视收益的完整分布,已在多个领域取得实证成功。DRL领域的核心任务之一是策略分布评估,即估计给定策略 $\pi$ 的收益分布 $\eta^\pi$。作为经典强化学习中时序差分算法的扩展,分布时序差分算法被提出。在表格情形下,\citet{rowland2018analysis} 与 \citet{rowland2023analysis} 分别证明了分布时序差分的两个实例——分类时序差分算法(CTD)和分位数时序差分算法(QTD)的渐近收敛性。本文进一步分析分布时序差分的有限样本性能。为便于理论分析,我们提出非参数分布时序差分算法(NTD)。针对具有状态空间 $S$ 和动作空间 $A$ 的 $\gamma$ 折扣无限时域表格型马尔可夫决策过程,我们证明:当通过 $p$-Wasserstein 距离度量估计误差时,NTD 算法需 $\wtilde O\prn{\frac{1}{\varepsilon^{2p}(1-\gamma)^{2p+2}}}$ 次迭代即可高概率达到 $\varepsilon$-最优估计量;在温和假设下,仅需 $\wtilde O\prn{\frac{1}{\varepsilon^{2}(1-\gamma)^{4}}}$ 次迭代即可保证 NTD 估计量 $\hat\eta^\pi$ 与 $\eta^\pi$ 的 Kolmogorov-Smirnov 距离以高概率小于 $\varepsilon$。我们重新审视 CTD 算法,证明在 $p$-Wasserstein 距离情形下,CTD 具有相同的非渐近收敛界。