We propose a novel visual SLAM method that integrates text objects tightly by treating them as semantic features via fully exploring their geometric and semantic prior. The text object is modeled as a texture-rich planar patch whose semantic meaning is extracted and updated on the fly for better data association. With the full exploration of locally planar characteristics and semantic meaning of text objects, the SLAM system becomes more accurate and robust even under challenging conditions such as image blurring, large viewpoint changes, and significant illumination variations (day and night). We tested our method in various scenes with the ground truth data. The results show that integrating texture features leads to a more superior SLAM system that can match images across day and night. The reconstructed semantic 3D text map could be useful for navigation and scene understanding in robotic and mixed reality applications. Our project page: https://github.com/SJTU-ViSYS/TextSLAM .
翻译:本文提出一种新颖的视觉SLAM方法,通过充分挖掘文本对象的几何与语义先验,将其作为语义特征紧密集成。文本对象被建模为纹理丰富的平面块,其语义信息在运行过程中被实时提取和更新,以提升数据关联的准确性。通过全面利用文本对象的局部平面特性与语义含义,SLAM系统在图像模糊、视角大幅变化及光照显著变化(昼夜交替)等挑战性条件下仍能保持更高的精度与鲁棒性。我们在多种场景中基于真实数据进行了方法测试,结果表明:集成纹理特征可构建更优越的SLAM系统,实现昼夜图像匹配。重建的语义三维文本地图可服务于机器人导航与混合现实应用中的场景理解。项目主页:https://github.com/SJTU-ViSYS/TextSLAM。