Natural policy gradient (NPG) and its variants are widely-used policy search methods in reinforcement learning. Inspired by prior work, a new NPG variant coined NPG-HM is developed in this paper, which utilizes the Hessian-aided momentum technique for variance reduction, while the sub-problem is solved via the stochastic gradient descent method. It is shown that NPG-HM can achieve the global last iterate $\epsilon$-optimality with a sample complexity of $\mathcal{O}(\epsilon^{-2})$, which is the best known result for natural policy gradient type methods under the generic Fisher non-degenerate policy parameterizations. The convergence analysis is built upon a relaxed weak gradient dominance property tailored for NPG under the compatible function approximation framework, as well as a neat way to decompose the error when handling the sub-problem. Moreover, numerical experiments on Mujoco-based environments demonstrate the superior performance of NPG-HM over other state-of-the-art policy gradient methods.
翻译:自然策略梯度及其变体是强化学习中广泛使用的策略搜索方法。受前期工作启发,本文提出了一种新的自然策略梯度变体NPG-HM,该方法利用Hessian辅助动量技术进行方差缩减,并通过随机梯度下降法求解子问题。研究表明,NPG-HM在样本复杂度为$\mathcal{O}(\epsilon^{-2})$的条件下可实现全局最后迭代$\epsilon$-最优性,这是通用Fisher非退化策略参数化下自然策略梯度类方法已知的最优结果。收敛分析建立在适配自然策略梯度的松弛弱梯度主导性质(基于兼容函数逼近框架)以及处理子问题时简洁的误差分解方式之上。此外,基于Mujoco环境的数值实验表明,NPG-HM的性能优于其他最先进的策略梯度方法。