Planning multi-contact motions in a receding horizon fashion requires a value function to guide the planning with respect to the future, e.g., building momentum to traverse large obstacles. Traditionally, the value function is approximated by computing trajectories in a prediction horizon (never executed) that foresees the future beyond the execution horizon. However, given the non-convex dynamics of multi-contact motions, this approach is computationally expensive. To enable online Receding Horizon Planning (RHP) of multi-contact motions, we find efficient approximations of the value function. Specifically, we propose a trajectory-based and a learning-based approach. In the former, namely RHP with Multiple Levels of Model Fidelity, we approximate the value function by computing the prediction horizon with a convex relaxed model. In the latter, namely Locally-Guided RHP, we learn an oracle to predict local objectives for locomotion tasks, and we use these local objectives to construct local value functions for guiding a short-horizon RHP. We evaluate both approaches in simulation by planning centroidal trajectories of a humanoid robot walking on moderate slopes, and on large slopes where the robot cannot maintain static balance. Our results show that locally-guided RHP achieves the best computation efficiency (95\%-98.6\% cycles converge online). This computation advantage enables us to demonstrate online receding horizon planning of our real-world humanoid robot Talos walking in dynamic environments that change on-the-fly.
翻译:在多接触运动进行递推水平规划时,需要借助值函数来引导对未来状态的决策(例如积累动量以跨越大型障碍物)。传统方法通过计算预测水平线内的轨迹(不执行)来近似值函数,以预见超越执行水平线的未来情况。然而,由于多接触运动的非凸动力学特性,该方法计算成本高昂。为实现多接触运动的在线递推水平规划,我们寻求值函数的高效近似方案。具体而言,我们提出了基于轨迹和基于学习的两种方法:前一种方法称为“多保真度模型递推水平规划”,通过凸松弛模型计算预测水平线来近似值函数;后一种方法称为“局部引导递推水平规划”,通过训练一个预测器来学习局部目标,并利用这些目标构建局部值函数以引导短时域递推水平规划。我们通过仿真评估了两种方法:在中等坡度步行及机器人无法维持静态平衡的大坡度场景下,规划人形机器人的质心轨迹。结果表明,局部引导递推水平规划实现了最优的计算效率(95%-98.6%的周期在线收敛)。这一计算优势使我们能够演示真实人形机器人Talos在动态变化环境中进行在线递推水平规划的能力。