Our work focuses on the Multi-Object Navigation (MultiON) task, where an agent needs to navigate to multiple objects in a given sequence. We systematically investigate the inherent modularity of this task by dividing our approach to contain four modules: (a) an object detection module trained to identify objects from RGB images, (b) a map building module to build a semantic map of the observed objects, (c) an exploration module enabling the agent to explore its surroundings, and finally (d) a navigation module to move to identified target objects. We focus on the navigation and the exploration modules in this work. We show that we can effectively leverage a PointGoal navigation model in the MultiON task instead of learning to navigate from scratch. Our experiments show that a PointGoal agent-based navigation module outperforms analytical path planning on the MultiON task. We also compare exploration strategies and surprisingly find that a random exploration strategy significantly outperforms more advanced exploration methods. We additionally create MultiON 2.0, a new large-scale dataset as a test-bed for our approach.
翻译:我们的工作聚焦于多目标导航(MultiON)任务,在该任务中,智能体需要按照给定顺序导航至多个目标。我们通过将方法划分为四个模块来系统性地研究该任务的内在模块化特性:(a)目标检测模块,用于从RGB图像中识别目标;(b)地图构建模块,用于构建观测目标的语义地图;(c)探索模块,使智能体能够探索其周围环境;以及(d)导航模块,用于移动至已识别的目标物体。本工作中,我们主要关注导航模块和探索模块。我们证明了在MultiON任务中可以有效利用PointGoal导航模型,而无需从头学习导航。实验表明,基于PointGoal智能体的导航模块在MultiON任务上的表现优于分析式路径规划方法。我们还比较了不同的探索策略,并意外发现随机探索策略在性能上显著优于更高级的探索方法。此外,我们创建了MultiON 2.0,一个大规模新数据集,作为我们方法的测试平台。