The prevailing paradigm for image-goal visual navigation often assumes access to large-scale datasets, substantial pretraining, and significant computational resources. In this work, we challenge this assumption. We show that we can collect a dataset, train an in-domain policy, and deploy it to the real world (1) in less than 120 minutes, (2) on a consumer laptop, (3) without any human intervention. Our method, MINav, formulates image-goal navigation as an offline goal-conditioned reinforcement learning problem, combining unsupervised data collection with hindsight goal relabeling and offline policy learning. Experiments in simulation and the real world show that MINav improves exploration efficiency, outperforms zero-shot navigation baselines in target environments, and scales favorably with dataset size. These results suggest that effective real-world robotic learning can be achieved with high computational efficiency, lowering the barrier to rapid policy prototyping and deployment.
翻译:图像目标视觉导航的主流范式通常假设可获取大规模数据集、充足的预训练和显著的计算资源。在这项工作中,我们挑战了这一假设。我们证明可以在(1)120分钟内、(2)在消费级笔记本电脑上、(3)无需任何人工干预的情况下,完成数据集收集、领域内策略训练和真实世界部署。我们的方法MINav将图像目标导航建模为离线目标条件强化学习问题,结合了无监督数据收集、事后目标重标记和离线策略学习。仿真与真实世界实验表明,MINav提升了探索效率,在目标环境中优于零样本导航基线,并展现出随数据集规模增加的良好扩展性。这些结果表明,高效计算能力即可实现有效的真实世界机器人学习,从而降低快速策略原型开发与部署的门槛。