This paper develops a policy learning method for tuning a pre-trained policy to adapt to additional tasks without altering the original task. A method named Adaptive Policy Gradient (APG) is proposed in this paper, which combines Bellman's principle of optimality with the policy gradient approach to improve the convergence rate. This paper provides theoretical analysis which guarantees the convergence rate and sample complexity of $\mathcal{O}(1/T)$ and $\mathcal{O}(1/\epsilon)$, respectively, where $T$ denotes the number of iterations and $\epsilon$ denotes the accuracy of the resulting stationary policy. Furthermore, several challenging numerical simulations, including cartpole, lunar lander, and robot arm, are provided to show that APG obtains similar performance compared to existing deterministic policy gradient methods while utilizing much less data and converging at a faster rate.
翻译:本文提出一种针对预训练策略的策略学习方法,使其在不改变原始任务的前提下自适应适配附加任务。文中提出名为自适应策略梯度(APG)的方法,结合贝尔曼最优性原理与策略梯度方法以提升收敛速度。本文提供理论分析,保证收敛速率和样本复杂度分别为$\mathcal{O}(1/T)$和$\mathcal{O}(1/\epsilon)$,其中$T$表示迭代次数,$\epsilon$表示所得稳态策略的精度。此外,通过包括小车爬坡、月球着陆器和机械臂在内的多项具有挑战性的数值仿真实验表明,APG在利用更少数据且以更快速率收敛的同时,可获得与现有确定性策略梯度方法相当的性能。