Images are a convenient way to specify which particular object instance an embodied agent should navigate to. Solving this task requires semantic visual reasoning and exploration of unknown environments. We present a system that can perform this task in both simulation and the real world. Our modular method solves sub-tasks of exploration, goal instance re-identification, goal localization, and local navigation. We re-identify the goal instance in egocentric vision using feature-matching and localize the goal instance by projecting matched features to a map. Each sub-task is solved using off-the-shelf components requiring zero fine-tuning. On the HM3D InstanceImageNav benchmark, this system outperforms a baseline end-to-end RL policy 7x and a state-of-the-art ImageNav model 2.3x (56% vs 25% success). We deploy this system to a mobile robot platform and demonstrate effective real-world performance, achieving an 88% success rate across a home and an office environment.
翻译:图像是一种便捷的方式,用于指定具身智能体应导航至的特定目标实例。解决该任务需要语义视觉推理以及对未知环境的探索。我们提出了一套可在仿真与真实世界中执行此任务的系统。本模块化方法解决了探索、目标实例重识别、目标定位及局部导航等子任务。我们通过特征匹配在自我中心视觉中重识别目标实例,并将匹配特征投影至地图以定位目标实例。每个子任务均使用无需微调的现成组件完成。在HM3D InstanceImageNav基准测试中,本系统性能超越基线端到端强化学习策略7倍,并领先最先进的ImageNav模型2.3倍(成功率56% vs 25%)。我们将该系统部署至移动机器人平台,在家庭与办公室环境中实现了88%的成功率,验证了其出色的真实世界性能。