Learning-based visual navigation for legged robots typically relies on continuous goal updates from hierarchical state estimation to provide a persistent directional reference. This reliance incurs additional sensory and computational overhead and deviates from fully end-to-end mobile autonomy. Furthermore, under partial observability, policies are prone to learn myopic behaviors, easily becoming trapped in dead ends and complex structural layouts. To address these limitations, we investigate a goal-initialized navigation setting, where the target is provided only once at the beginning of an episode, requiring the robot to operate based on intrinsic spatial memory without subsequent goal updates from external modules. In this work, we propose GUIDE, a fully end-to-end reinforcement learning framework designed to cultivate internal directional awareness. Specifically, GUIDE incorporates a spatial anchor predictor that leverages multi-frequency proprioceptive history to extract egomotion representations, thereby maintaining a persistent long-horizon spatial context for navigation. Concurrently, it utilizes raw depth streams to perceive local environmental geometry. We evaluate the proposed framework across both simulation and real-world scenarios on a quadruped robot. Experiments show that GUIDE learns reliable egomotion and directional awareness, enabling a fully end-to-end deployed policy to safely navigate through dense clutter and structured mazes without subsequent goal guidance or prior maps.
翻译:基于学习的腿式机器人视觉导航通常依赖层级状态估计的连续目标更新来提供持久的方向参考。这种依赖性会引入额外的传感与计算开销,并偏离全自主端到端移动能力。此外,在部分可观测条件下,策略容易学习短视行为,陷入死胡同和复杂结构布局。为解决这些局限,我们研究了一种目标初始化导航场景:目标仅在回合开始时一次性提供,要求机器人基于内在空间记忆运行,无需依赖外部模块的后续目标更新。本文提出GUIDE——一种全端到端强化学习框架,旨在培养内在方向感知能力。具体而言,GUIDE包含空间锚点预测器,它利用多频率本体感受历史提取自运动表征,从而维持导航所需的持久长程空间上下文;同时,通过原始深度流感知局部环境几何结构。我们在四足机器人的仿真与真实场景中评估了所提框架。实验表明,GUIDE能学习可靠的自运动与方向感知能力,使全端到端部署的策略无需后续目标引导或先验地图即可安全穿越密集杂乱区域与结构化迷宫。