Despite impressive results, reinforcement learning (RL) suffers from slow convergence and requires a large variety of tuning strategies. In this paper, we investigate the ability of RL algorithms on simple continuous control tasks. We show that without reward and environment tuning, RL suffers from poor convergence. In turn, we introduce an optimal control (OC) theoretic learning-based method that can solve the same problems robustly with simple parsimonious costs. We use the Hamilton-Jacobi-Bellman (HJB) and first-order gradients to learn optimal time-varying value functions and therefore, policies. We show the relaxation of our objective results in time-varying Lyapunov functions, further verifying our approach by providing guarantees over a compact set of initial conditions. We compare our method to Soft Actor Critic (SAC) and Proximal Policy Optimisation (PPO). In this comparison, we solve all tasks, we never underperform in task cost and we show that at the point of our convergence, we outperform SAC and PPO in the best case by 4 and 2 orders of magnitude.
翻译:尽管强化学习取得了显著成果,但其收敛速度缓慢且需要大量调参策略。本文研究了强化学习算法在简单连续控制任务中的表现。研究表明,在没有奖励和环境调参的情况下,强化学习存在收敛性差的问题。为此,我们提出一种基于最优控制(OC)理论的学习方法,该方法能够通过简单的简约代价函数稳健地解决相同问题。我们利用哈密顿-雅可比-贝尔曼(HJB)方程和一阶梯度学习最优时变值函数及其相应的策略。我们证明了目标函数的松弛化会得到时变李雅普诺夫函数,并通过在一组紧凑的初始条件上提供保证进一步验证了该方法。我们将所提方法与软演员-评论家(SAC)和近端策略优化(PPO)进行了比较。在比较中,我们解决了所有任务,从未在任务代价上表现不佳,并且表明在收敛点上,我们在最优情况下分别比SAC和PPO高出4个和2个数量级。