A fundamental task in robotics is to navigate between two locations. In particular, real-world navigation can require long-horizon planning using high-dimensional RGB images, which poses a substantial challenge for end-to-end learning-based approaches. Current semi-parametric methods instead achieve long-horizon navigation by combining learned modules with a topological memory of the environment, often represented as a graph over previously collected images. However, using these graphs in practice requires tuning a number of pruning heuristics. These heuristics are necessary to avoid spurious edges, limit runtime memory usage and maintain reasonably fast graph queries in large environments. In this work, we present One-4-All (O4A), a method leveraging self-supervised and manifold learning to obtain a graph-free, end-to-end navigation pipeline in which the goal is specified as an image. Navigation is achieved by greedily minimizing a potential function defined continuously over image embeddings. Our system is trained offline on non-expert exploration sequences of RGB data and controls, and does not require any depth or pose measurements. We show that O4A can reach long-range goals in 8 simulated Gibson indoor environments and that resulting embeddings are topologically similar to ground truth maps, even if no pose is observed. We further demonstrate successful real-world navigation using a Jackal UGV platform.
翻译:机器人学中的一项基本任务是在两个位置之间导航。特别是在现实世界中,导航可能需要利用高维RGB图像进行长时域规划,这对基于端到端学习的方法构成了重大挑战。当前的半参数化方法通过将学习模块与环境拓扑记忆(通常表示为基于先前采集图像的图结构)相结合,来实现长时域导航。然而,在实际使用这些图结构时,需要调整大量剪枝启发式规则。这些启发式规则对于避免虚假边、限制运行时内存消耗以及在大规模环境中维持合理的图查询速度至关重要。在本工作中,我们提出One-4-All (O4A)方法,该方法利用自监督学习和流形学习,构建了一种无图结构的端到端导航流水线,其中目标以图像形式指定。导航通过贪婪最小化在图像嵌入空间连续定义的势函数来实现。我们的系统基于非专家探索序列的RGB数据与控制指令进行离线训练,且无需任何深度或位姿测量。实验表明,O4A可在8个模拟Gibson室内环境中实现长距离目标到达,且生成的嵌入在拓扑结构上与真实地图高度相似,即使未观测到位姿信息。此外,我们通过Jackal UGV平台验证了其在真实世界中的成功导航能力。