Geometric navigation is nowadays a well-established field of robotics and the research focus is shifting towards higher-level scene understanding, such as Semantic Mapping. When a robot needs to interact with its environment, it must be able to comprehend the contextual information of its surroundings. This work focuses on classifying and localising objects within a map, which is under construction (SLAM) or already built. To further explore this direction, we propose a framework that can autonomously detect and localize predefined objects in a known environment using a multi-modal sensor fusion approach (combining RGB and depth data from an RGB-D camera and a lidar). The framework consists of three key elements: understanding the environment through RGB data, estimating depth through multi-modal sensor fusion, and managing artifacts (i.e., filtering and stabilizing measurements). The experiments show that the proposed framework can accurately detect 98% of the objects in the real sample environment, without post-processing, while 85% and 80% of the objects were mapped using the single RGBD camera or RGB + lidar setup respectively. The comparison with single-sensor (camera or lidar) experiments is performed to show that sensor fusion allows the robot to accurately detect near and far obstacles, which would have been noisy or imprecise in a purely visual or laser-based approach.
翻译:几何导航如今已是机器人领域成熟的研究方向,研究重心正转向更高层次的场景理解,例如语义映射。当机器人需要与环境交互时,必须能够理解周围环境的上下文信息。本研究聚焦于在建图(SLAM)或已有地图中分类与定位目标。为深入探索这一方向,我们提出了一种基于多模态传感器融合(结合RGB-D相机的RGB与深度数据及激光雷达)的框架,可在已知环境中自主检测并定位预设目标。该框架包含三个关键环节:通过RGB数据理解环境、通过多模态传感器融合估算深度,以及管理人工制品(即滤波与稳定测量值)。实验表明,该框架无需后处理即可在真实样本环境中准确检测98%的目标;而仅使用单个RGBD相机或RGB+激光雷达配置时,分别有85%和80%的目标被成功映射。通过与单传感器(相机或激光雷达)实验对比可知,传感器融合使机器人能够准确检测远近障碍物——纯视觉或纯激光方法对此类障碍物的检测结果往往存在噪声或不精确的问题。