Deep reinforcement learning (RL) has been endowed with high expectations in tackling challenging manipulation tasks in an autonomous and self-directed fashion. Despite the significant strides made in the development of reinforcement learning, the practical deployment of this paradigm is hindered by at least two barriers, namely, the engineering of a reward function and ensuring the safety guaranty of learning-based controllers. In this paper, we address these challenging limitations by proposing a framework that merges a reinforcement learning \lstinline[columns=fixed]{planner} that is trained using sparse rewards with a model predictive controller (MPC) \lstinline[columns=fixed]{actor}, thereby offering a safe policy. On the one hand, the RL \lstinline[columns=fixed]{planner} learns from sparse rewards by selecting intermediate goals that are easy to achieve in the short term and promising to lead to target goals in the long term. On the other hand, the MPC \lstinline[columns=fixed]{actor} takes the suggested intermediate goals from the RL \lstinline[columns=fixed]{planner} as the input and predicts how the robot's action will enable it to reach that goal while avoiding any obstacles over a short period of time. We evaluated our method on four challenging manipulation tasks with dynamic obstacles and the results demonstrate that, by leveraging the complementary strengths of these two components, the agent can solve manipulation tasks in complex, dynamic environments safely with a $100\%$ success rate. Videos are available at \url{https://videoviewsite.wixsite.com/mpc-hgg}.
翻译:深度强化学习(RL)在自主解决复杂操作任务方面被寄予厚望。尽管强化学习取得了显著进展,但其实践部署仍面临至少两大障碍:奖励函数的设计工程以及确保基于学习的控制器的安全性保证。本文通过提出一个融合稀疏奖励训练的强化学习规划器(planner)与模型预测控制器(MPC actor)的框架来解决这些挑战性限制,从而提供安全策略。一方面,RL规划器通过选择短期易于实现且有望长期达成目标中的中间目标,从稀疏奖励中学习。另一方面,MPC执行器将RL规划器建议的中间目标作为输入,预测机器人动作如何使其在短时间内避开障碍物并达成该目标。我们在四个包含动态障碍物的挑战性操作任务上评估了该方法,结果表明,通过利用这两个组件的互补优势,智能体能够在复杂动态环境中以100%的成功率安全完成操作任务。视频地址:\url{https://videoviewsite.wixsite.com/mpc-hgg}