Re-localizing a camera from a single image in a previously mapped area is vital for many computer vision applications in robotics and augmented/virtual reality. In this work, we address the problem of estimating the 6 DoF camera pose relative to a global frame from a single image. We propose to leverage a novel network of relative spatial and temporal geometric constraints to guide the training of a Deep Network for localization. We employ simultaneously spatial and temporal relative pose constraints that are obtained not only from adjacent camera frames but also from camera frames that are distant in the spatio-temporal space of the scene. We show that our method, through these constraints, is capable of learning to localize when little or very sparse ground-truth 3D coordinates are available. In our experiments, this is less than 1% of available ground-truth data. We evaluate our method on 3 common visual localization datasets and show that it outperforms other direct pose estimation methods.
翻译:在先前建图区域中通过单幅图像重新定位摄像机,对机器人学和增强/虚拟现实领域的诸多计算机视觉应用至关重要。本文研究从单幅图像估计相对于全局坐标系的六自由度摄像机位姿问题,提出利用新型相对空间与时间几何约束网络来指导深度网络定位训练。我们同步使用通过相邻摄像机帧以及场景时空空间中相距较远的摄像机帧获取的相对空间与时间位姿约束。实验表明,该方法在仅有少量甚至极其稀疏的真实三维坐标可用时(本实验中可用真实数据不足总量的1%),仍能通过此类约束习得定位能力。我们在三个常用视觉定位数据集上评估该方法,结果显示其性能优于其他直接位姿估计方法。