As a new embodied vision task, Instance ImageGoal Navigation (IIN) aims to navigate to a specified object depicted by a goal image in an unexplored environment. The main challenge of this task lies in identifying the target object from different viewpoints while rejecting similar distractors. Existing ImageGoal Navigation methods usually adopt the simple Exploration-Exploitation framework and ignore the identification of specific instance during navigation. In this work, we propose to imitate the human behaviour of ``getting closer to confirm" when distinguishing objects from a distance. Specifically, we design a new modular navigation framework named Instance-aware Exploration-Verification-Exploitation (IEVE) for instance-level image goal navigation. Our method allows for active switching among the exploration, verification, and exploitation actions, thereby facilitating the agent in making reasonable decisions under different situations. On the challenging HabitatMatterport 3D semantic (HM3D-SEM) dataset, our method surpasses previous state-of-the-art work, with a classical segmentation model (0.684 vs. 0.561 success) or a robust model (0.702 vs. 0.561 success)
翻译:作为一种新的具身视觉任务,实例图像目标导航(Instance ImageGoal Navigation, IIN)旨在通过目标图像描述的指定物体,在未探索环境中进行导航。该任务的主要挑战在于从不同视角识别目标物体,同时排除相似的干扰物。现有的图像目标导航方法通常采用简单的探索-利用框架,而忽略了导航过程中对特定实例的识别。在本工作中,我们提出模仿人类在远距离区分物体时“靠近确认”的行为。具体而言,我们设计了一种名为实例感知的探索-验证-利用(Instance-aware Exploration-Verification-Exploitation, IEVE)的新型模块化导航框架,用于实例级别的图像目标导航。该方法支持在探索、验证和利用动作之间主动切换,从而使智能体在不同情况下做出合理决策。在具有挑战性的HabitatMatterport 3D语义(HM3D-SEM)数据集上,我们的方法超越了先前的最优工作,无论是使用经典分割模型(成功率0.684 vs. 0.561)还是鲁棒模型(成功率0.702 vs. 0.561)。