Multi-camera systems have been shown to improve the accuracy and robustness of SLAM estimates, yet state-of-the-art SLAM systems predominantly support monocular or stereo setups. This paper presents a generic sparse visual SLAM framework capable of running on any number of cameras and in any arrangement. Our SLAM system uses the generalized camera model, which allows us to represent an arbitrary multi-camera system as a single imaging device. Additionally, it takes advantage of the overlapping fields of view (FoV) by extracting cross-matched features across cameras in the rig. This limits the linear rise in the number of features with the number of cameras and keeps the computational load in check while enabling an accurate representation of the scene. We evaluate our method in terms of accuracy, robustness, and run time on indoor and outdoor datasets that include challenging real-world scenarios such as narrow corridors, featureless spaces, and dynamic objects. We show that our system can adapt to different camera configurations and allows real-time execution for typical robotic applications. Finally, we benchmark the impact of the critical design parameters - the number of cameras and the overlap between their FoV that define the camera configuration for SLAM. All our software and datasets are freely available for further research.
翻译:多相机系统已被证明能够提升SLAM估计的精度与鲁棒性,然而当前最先进的SLAM系统主要支持单目或立体视觉配置。本文提出一种通用的稀疏视觉SLAM框架,能够适配任意数量及任意排列的相机。本SLAM系统采用广义相机模型,可将任意多相机系统表示为单一成像设备。此外,系统通过提取相机阵列中的跨相机交叉匹配特征,充分利用视场重叠区域,从而限制特征数量随相机数量线性增长,在保持场景精确表征的同时控制计算负载。我们在包含狭窄走廊、无纹理区域及动态物体等挑战性真实场景的室内外数据集上,从精度、鲁棒性及运行时间三个维度对方法进行评估。实验表明,本系统可适配不同相机配置,并在典型机器人应用中实现实时运行。最后,我们基准测试了关键设计参数——相机数量及定义SLAM相机配置的视场重叠程度——对系统性能的影响。所有软件及数据集均已开源以供后续研究。