We present a real-time visual-inertial dense mapping method capable of performing incremental 3D mesh reconstruction with high quality using only sequential monocular images and inertial measurement unit (IMU) readings. 6-DoF camera poses are estimated by a robust feature-based visual-inertial odometry (VIO), which also generates noisy sparse 3D map points as a by-product. We propose a sparse point aided multi-view stereo neural network (SPA-MVSNet) that can effectively leverage the informative but noisy sparse points from the VIO system. The sparse depth from VIO is firstly completed by a single-view depth completion network. This dense depth map, although naturally limited in accuracy, is then used as a prior to guide our MVS network in the cost volume generation and regularization for accurate dense depth prediction. Predicted depth maps of keyframe images by the MVS network are incrementally fused into a global map using TSDF-Fusion. We extensively evaluate both the proposed SPA-MVSNet and the entire visual-inertial dense mapping system on several public datasets as well as our own dataset, demonstrating the system's impressive generalization capabilities and its ability to deliver high-quality 3D mesh reconstruction online. Our proposed dense mapping system achieves a 39.7% improvement in F-score over existing systems when evaluated on the challenging scenarios of the EuRoC dataset.
翻译:我们提出了一种实时视觉-惯性密集建图方法,该方法仅利用顺序单目图像和惯性测量单元(IMU)读数,即可实现高质量的三维网格增量式重建。六自由度相机位姿通过基于稳健特征的视觉-惯性里程计(VIO)进行估计,该过程同时生成带噪声的稀疏三维地图点作为副产品。我们提出了一种稀疏点辅助的多视角立体神经网络(SPA-MVSNet),能够有效利用VIO系统中包含信息但含噪的稀疏点。首先通过单视角深度补全网络对VIO输出的稀疏深度进行补全。尽管该稠密深度图在精度上天然受限,但其随后作为先验信息,用于引导我们的MVS网络在代价体生成和正则化过程中实现精确稠密深度预测。MVS网络对关键帧图像预测的深度图通过TSDF融合算法增量式融合至全局地图中。我们在多个公开数据集及自有数据集上对提出的SPA-MVSNet及整个视觉-惯性密集建图系统进行了全面评估,验证了系统出色的泛化能力及在线生成高质量三维网格重建的能力。在EuRoC数据集的挑战性场景上,我们的密集建图系统相较于现有系统实现了F-score 39.7%的提升。