The assumption of a static environment is common in many geometric computer vision tasks like SLAM but limits their applicability in highly dynamic scenes. Since these tasks rely on identifying point correspondences between input images within the static part of the environment, we propose a graph neural network-based sparse feature matching network designed to perform robust matching under challenging conditions while excluding keypoints on moving objects. We employ a similar scheme of attentional aggregation over graph edges to enhance keypoint representations as state-of-the-art feature-matching networks but augment the graph with epipolar and temporal information and vastly reduce the number of graph edges. Furthermore, we introduce a self-supervised training scheme to extract pseudo labels for image pairs in dynamic environments from exclusively unprocessed visual-inertial data. A series of experiments show the superior performance of our network as it excludes keypoints on moving objects compared to state-of-the-art feature matching networks while still achieving similar results regarding conventional matching metrics. When integrated into a SLAM system, our network significantly improves performance, especially in highly dynamic scenes.
翻译:静态环境假设是SLAM等几何计算机视觉任务的常见前提,但严重制约了其在高度动态场景中的适用性。由于此类任务依赖于在环境静态部分识别输入图像间的点对应关系,我们提出一种基于图神经网络的稀疏特征匹配网络,旨在挑战性条件下实现鲁棒匹配,同时排除运动物体上的关键点。本文采用与现有最优特征匹配网络类似的跨图边注意力聚合机制增强关键点表征,但通过补充极几何与时间信息对图结构进行增强,并大幅减少图边的数量。此外,我们提出一种自监督训练方案,仅利用未处理的视觉惯性数据即可为动态环境中的图像对提取伪标签。系列实验表明,该网络在排除运动物体关键点的同时,仍能在常规匹配指标上达到与现有最优特征匹配网络相当的性能。当集成至SLAM系统后,该网络显著提升了系统性能,尤其是在高度动态场景中。