We present a footstep planning policy for quadrupedal locomotion that is able to directly take into consideration a-priori safety information in its decisions. At its core, a learning process analyzes terrain patches, classifying each landing location by its kinematic feasibility, shin collision, and terrain roughness. This information is then encoded into a small vector representation and passed as an additional state to the footstep planning policy, which furthermore proposes only safe footstep location by applying a masked variant of the Proximal Policy Optimization algorithm. The performance of the proposed approach is shown by comparative simulations and experiments on an electric quadruped robot walking in different rough terrain scenarios. We show that violations of the above safety conditions are greatly reduced both during training and the successive deployment of the policy, resulting in an inherently safer footstep planner. Furthermore, we show how, as a byproduct, fewer reward terms are needed to shape the behavior of the policy, which in return is able to achieve both better final performances and sample efficiency.
翻译:我们提出了一种四足机器人落脚规划策略,该策略能够直接在其决策中考虑先验安全信息。其核心是一个学习过程,通过分析地形块对每个落足位置的运动学可行性、小腿碰撞风险和地形粗糙度进行分类。这些信息随后被编码成一个小型向量表示,作为额外状态传递给落脚规划策略。该策略通过应用掩码化近端策略优化算法,仅提出安全的落脚位置。通过在不同崎岖地形场景下与电动四足机器人的对比仿真和实验,展示了所提方法的性能。结果表明,在策略训练和后续部署过程中,上述安全条件的违反率显著降低,从而形成了本质更安全的落脚规划器。此外,作为副产品,塑造策略行为所需的奖励项更少,这使得策略能够同时获得更好的最终性能和样本效率。