Vision-based navigation models, particularly foundation models, generate viable trajectories from RGB observations alone. However, even state-of-the-art transformer- and diffusion-based policies struggle to generalize in unfamiliar deployment environments containing unseen obstacles or shifted conditions. The resulting trajectories often remain goal-directed but unsafe. Existing efforts improve safety through external trajectory correction or internal geometric priors, yet the resulting policies are not trained to explicitly represent obstacle boundaries or traversable free-space structure. To address this, we propose a navigation model that incorporates these structures directly into the policy via fine-tuning and is designed to be compatible with diverse RGB-based backbones. Across multiple robot platforms, indoor environments, and static and dynamic obstacle scenarios, our method reduces collision frequency relative to ViNT, NoMaD, and their CARE-augmented variants while maintaining goal-reaching performance.
翻译:基于视觉的导航模型,特别是基础模型,能够仅从RGB观测生成可行轨迹。然而,即使是最先进的基于Transformer和扩散的策略,也难以在包含未见障碍物或环境条件偏移的陌生部署场景中泛化。由此产生的轨迹通常仍以目标为导向,但存在安全隐患。现有研究通过外部轨迹校正或内部几何先验来提升安全性,但所得策略并未被训练以显式表征障碍物边界或可通行自由空间结构。为解决这一问题,我们提出一种导航模型,该模型通过微调将上述结构直接融入策略,并设计为兼容多种基于RGB的骨干网络。在多种机器人平台、室内环境以及静态与动态障碍物场景中,相较于ViNT、NoMaD及其经CARE增强的变体,我们的方法在保持目标到达性能的同时降低了碰撞频率。