Visual localization is a core component in many applications, including augmented reality (AR). Localization algorithms compute the camera pose of a query image w.r.t. a scene representation, which is typically built from images. This often requires capturing and storing large amounts of data, followed by running Structure-from-Motion (SfM) algorithms. An interesting, and underexplored, source of data for building scene representations are 3D models that are readily available on the Internet, e.g., hand-drawn CAD models, 3D models generated from building footprints, or from aerial images. These models allow to perform visual localization right away without the time-consuming scene capturing and model building steps. Yet, it also comes with challenges as the available 3D models are often imperfect reflections of reality. E.g., the models might only have generic or no textures at all, might only provide a simple approximation of the scene geometry, or might be stretched. This paper studies how the imperfections of these models affect localization accuracy. We create a new benchmark for this task and provide a detailed experimental evaluation based on multiple 3D models per scene. We show that 3D models from the Internet show promise as an easy-to-obtain scene representation. At the same time, there is significant room for improvement for visual localization pipelines. To foster research on this interesting and challenging task, we release our benchmark at v-pnk.github.io/cadloc.
翻译:视觉定位是许多应用(包括增强现实)中的核心组成部分。定位算法计算查询图像相对于场景表示的相机位姿,该场景表示通常由图像构建。这通常需要捕获和存储大量数据,随后运行运动恢复结构算法。一个有趣且尚未充分探索的用于构建场景表示的数据来源是互联网上现成的三维模型,例如手绘CAD模型、基于建筑轮廓生成的三维模型或来自航拍图像的三维模型。这些模型允许直接进行视觉定位,无需耗时的场景捕获和模型构建步骤。然而,这也带来了挑战,因为可用的三维模型往往是对现实的不完美反映。例如,模型可能仅有通用纹理或无纹理,可能仅提供场景几何的简单近似,或可能存在拉伸变形。本文研究了这些模型的不完美性如何影响定位精度。我们为此任务创建了一个新基准,并基于每个场景的多个三维模型提供了详细的实验评估。我们表明,来自互联网的三维模型作为易于获取的场景表示具有潜力。同时,视觉定位流程仍有显著改进空间。为促进对这一有趣且富有挑战性任务的研究,我们在v-pnk.github.io/cadloc发布了我们的基准。