Learning-based visual navigation has enhanced semantic goal-reaching capabilities. However, due to their black-box nature, purely end-to-end models often lack explicit geometric constraints, leading to unpredictable and unreliable obstacle avoidance in open environments. Conversely, traditional geometric planners ensure safety but struggle with high-dimensional visual targets. To address these limitations, we propose SemGeoNav, a novel hierarchical visual navigation framework.It tightly integrates the high-level semantic reasoning of end-to-end models with the reliable local planning ability of geometry-based methods, achieving robust image-based navigation while significantly improving obstacle avoidance. Furthermore, we introduce a temporal trajectory smoothing mechanism to ensure continuous and stable robot motion. We evaluated SemGeoNav on a Unitree Go2 quadruped robot in real-world environments. The results demonstrate that SemGeoNav outperforms existing representative methods, including ViNT and NoMaD, achieving higher success rates and shorter navigation times.
翻译:基于学习的视觉导航已提升了语义目标到达能力。然而,纯端到端模型因其黑箱特性,常缺乏显式几何约束,导致开阔环境中避障行为的不可预测性与不可靠性。相比之下,传统几何规划器能保障安全,却难以处理高维视觉目标。为克服上述局限,我们提出SemGeoNav——一种新颖的分层视觉导航框架。该框架紧密融合端到端模型的高层语义推理能力与基于几何方法的可靠局部规划能力,在实现鲁棒图像导航的同时显著提升避障性能。此外,我们引入时序轨迹平滑机制以确保机器人运动的连续性与稳定性。在真实环境中基于Unitree Go2四足机器人平台开展的实验表明:SemGeoNav在成功率和导航耗时两方面均优于ViNT、NoMaD等现有代表性方法。