Enabling humanoid robots to operate in complex, dynamic environments remains a critical challenge, fundamentally limited by the ability to navigate robustly, safely, and accurately. While reinforcement learning with velocity-commanded policies has achieved remarkable robustness in humanoid locomotion, this approach lacks explicit control of the foothold placement, leading to unsafe behavior, such as stepping onto human feet, or imprecise navigation, hindering the following manipulation task. Conversely, explicit foothold-tracking policies offer a promising alternative by directly being commanded with target foot poses. However, existing approaches are often limited by unrealistic state assumptions, compromising real-world deployment, or they are part of staged pipelines, making them tied to specific downstream tasks. In this work, we introduce a novel, lightweight framework for training general-purpose 3D foothold-tracking policies. By dynamically providing footstep support through a goal sampler, this method enables the learned policy to be agnostic to specific terrains. Our new target representation effectively mitigates challenges arising in the real world, such as noisy and inaccurate pose estimation and foot contact estimation. Designed for direct real-world transfer, our policy acts as a standalone low-level controller that can be seamlessly paired with various high-level foothold generators. We demonstrate the effectiveness of our framework through extensive experiments in simulation and in the real world. By coupling our policy with different upstream planners, we achieve natural and accurate locomotion in challenging settings, paving the way for loco-manipulation tasks in complex environments.
翻译:使类人机器人在复杂动态环境中运行仍是一项关键挑战,其根本限制在于稳健、安全且精准的导航能力。尽管采用速度指令策略的强化学习在类人机器人运动方面取得了显著鲁棒性,但该方法缺乏对落脚点位置的显式控制,导致不安全行为(例如踩到人类脚部)或导航不精确,进而阻碍后续操作任务。相比之下,显式落脚点跟踪策略通过直接以目标足部位姿作为指令提供了一种有前景的替代方案。然而,现有方法常受限于不切实际的状态假设(影响实际部署),或作为分阶段流水线的一部分而与特定下游任务绑定。本研究提出了一种用于训练通用三维落脚点跟踪策略的新型轻量级框架。通过目标采样器动态提供落脚支撑,该方法使所学策略能够独立于特定地形。我们提出的新目标表征有效缓解了现实世界中出现的挑战,例如带有噪声且不准确的姿态估计与足部接触估计。该策略专为直接迁移至真实场景设计,可作为独立的底层控制器,与各类高层落脚点生成器无缝配合。通过仿真与真实世界中的大量实验,我们验证了该框架的有效性。将该策略与不同的上游规划器相结合,我们在具有挑战性的环境中实现了自然且精准的运动,为复杂环境下的机械臂-运动协同操作任务铺平了道路。