Understanding and mapping a new environment are core abilities of any autonomously navigating agent. While classical robotics usually estimates maps in a stand-alone manner with SLAM variants, which maintain a topological or metric representation, end-to-end learning of navigation keeps some form of memory in a neural network. Networks are typically imbued with inductive biases, which can range from vectorial representations to birds-eye metric tensors or topological structures. In this work, we propose to structure neural networks with two neural implicit representations, which are learned dynamically during each episode and map the content of the scene: (i) the Semantic Finder predicts the position of a previously seen queried object; (ii) the Occupancy and Exploration Implicit Representation encapsulates information about explored area and obstacles, and is queried with a novel global read mechanism which directly maps from function space to a usable embedding space. Both representations are leveraged by an agent trained with Reinforcement Learning (RL) and learned online during each episode. We evaluate the agent on Multi-Object Navigation and show the high impact of using neural implicit representations as a memory source.
翻译:理解和映射新环境是任何自主导航智能体的核心能力。经典机器人通常通过SLAM变体以独立方式估计地图,维护拓扑或度量表示;而端到端的导航学习则在神经网络中保留某种形式的记忆。网络通常被赋予归纳偏置,其范围可从向量化表示到鸟瞰度量张量或拓扑结构。在这项工作中,我们提出用两种神经隐式表示来结构化神经网络,这些表示在每轮交互中动态学习并映射场景内容:(i)语义查找器预测先前观察到的查询对象的位置;(ii)占据与探索隐式表示封装了已探索区域和障碍物的信息,并通过一种新颖的全局读取机制进行查询,该机制直接从函数空间映射到可用的嵌入空间。两种表示均由通过强化学习(RL)训练的智能体在每轮交互中在线学习利用。我们在多目标导航任务上评估了该智能体,并展示了使用神经隐式表示作为记忆源的显著影响。