Deep reinforcement learning (DRL) has emerged as a promising solution to mastering explosive and versatile quadrupedal jumping skills. However, current DRL-based frameworks usually rely on well-defined reference trajectories, which are obtained by capturing animal motions or transferring experience from existing controllers. This work explores the possibility of learning dynamic jumping without imitating a reference trajectory. To this end, we incorporate a curriculum design into DRL so as to accomplish challenging tasks progressively. Starting from a vertical in-place jump, we then generalize the learned policy to forward and diagonal jumps and, finally, learn to jump across obstacles. Conditioned on the desired landing location, orientation, and obstacle dimensions, the proposed approach contributes to a wide range of jumping motions, including omnidirectional jumping and robust jumping, alleviating the effort to extract references in advance. Particularly, without constraints from the reference motion, a 90cm forward jump is achieved, exceeding previous records for similar robots reported in the existing literature. Additionally, continuous jumping on the soft grassy floor is accomplished, even when it is not encountered in the training stage. A supplementary video showing our results can be found at https://youtu.be/nRaMCrwU5X8 .
翻译:深度强化学习(DRL)已成为掌握爆发性、多变的四足机器人跳跃技能的一种有前景的解决方案。然而,当前基于DRL的框架通常依赖于精确定义的参考轨迹——这些轨迹通过捕捉动物运动或从现有控制器中迁移经验获得。本研究探索了在不模仿参考轨迹的情况下学习动态跳跃的可能性。为此,我们将课程设计融入DRL,以逐步完成具有挑战性的任务。从原地垂直跳跃开始,我们将所学策略推广到前向和对角跳跃,最终学会跨越障碍。基于期望的着陆位置、朝向和障碍物尺寸,所提出的方法产生了包括全向跳跃和鲁棒跳跃在内的广泛跳跃动作,减轻了预先提取参考轨迹的负担。特别地,在无参考运动约束的情况下,实现了前向90厘米跳跃,超越了现有文献中类似机器人此前报告的记录。此外,即使训练阶段未涉及松软草地,机器人也成功完成了连续跳跃。展示我们结果的补充视频见https://youtu.be/nRaMCrwU5X8。