Recently, Segment Anything Model (SAM) has demonstrated strong generalizability in various instance segmentation tasks. However, its performance is severely dependent on the quality of manual prompts. In addition, the RGB images that instance segmentation methods normally use inherently lack depth information. As a result, the ability of these methods to perceive spatial structures and delineate object boundaries is hindered. To address these challenges, we propose a Self-prompted Depth-Aware SAM (SPDA-SAM) for instance segmentation. Specifically, we design a Semantic-Spatial Self-prompt Module (SSSPM) which extracts the semantic and spatial prompts from the image encoder and the mask decoder of SAM, respectively. Furthermore, we introduce a Coarse-to-Fine RGB-D Fusion Module (C2FFM), in which the features extracted from a monocular RGB image and the depth map estimated from it are fused. In particular, the structural information in the depth map is used to provide coarse-grained guidance to feature fusion, while local variations in depth are encoded in order to fuse fine-grained feature representations. To our knowledge, SAM has not been explored in such self-prompted and depth-aware manners. Experimental results demonstrate that our SPDA-SAM outperforms its state-of-the-art counterparts across twelve different data sets. These promising results should be due to the guidance of the self-prompts and the compensation for the spatial information loss by the coarse-to-fine RGB-D fusion operation.
翻译:最近,分割一切模型(Segment Anything Model, SAM)在各种实例分割任务中展现出强大的泛化能力。然而,其性能严重依赖于人工提示的质量。此外,实例分割方法通常使用的RGB图像本质上缺乏深度信息,这制约了这些方法感知空间结构和描绘物体边界的能力。为应对这些挑战,我们提出了一种用于实例分割的自提示深度感知SAM(SPDA-SAM)。具体而言,我们设计了语义-空间自提示模块(SSSPM),该模块分别从SAM的图像编码器和掩码解码器中提取语义和空间提示。此外,我们引入了一种从粗到细的RGB-D融合模块(C2FFM),其中,从单目RGB图像中提取的特征与由其估计出的深度图相融合。特别地,利用深度图中的结构信息为特征融合提供粗粒度引导,同时编码深度的局部变化以融合细粒度特征表示。据我们所知,SAM此前尚未以这种自提示和深度感知的方式被探索。实验结果表明,我们的SPDA-SAM在十二个不同数据集上均优于当前最先进的同类方法。这些令人瞩目的结果应归因于自提示的引导以及从粗到细的RGB-D融合操作对空间信息丢失的补偿。