In this work, we consider the complex control problem of making a monopod reach a target with a jump. The monopod can jump in any direction and the terrain underneath its foot can be uneven. This is a template of a much larger class of problems, which are extremely challenging and computationally expensive to solve using standard optimisation-based techniques. Reinforcement Learning (RL) could be an interesting alternative, but the application of an end-to-end approach in which the controller must learn everything from scratch, is impractical. The solution advocated in this paper is to guide the learning process within an RL framework by injecting physical knowledge. This expedient brings to widespread benefits, such as a drastic reduction of the learning time, and the ability to learn and compensate for possible errors in the low-level controller executing the motion. We demonstrate the advantage of our approach with respect to both optimization-based and end-to-end RL approaches.
翻译:本文研究单足机器人通过跳跃到达目标位置的复杂控制问题。该单足机器人能够向任意方向跳跃,且其足部下方的地形可能不平坦。这属于一类具有更广泛代表性的模板问题,采用传统基于优化的方法求解时计算成本极高且极具挑战性。强化学习(RL)可作为替代方案,但若采用端到端方法令控制器从零开始学习一切,则缺乏实际可行性。本文提出的解决方案是在强化学习框架中注入物理先验知识来引导学习过程。这一策略带来了广泛优势,包括大幅缩短学习时间,以及能够学习并补偿执行动作的低层控制器中的潜在误差。我们通过实验证明了本方法相较于基于优化的方法与端到端强化学习方法的优越性。