Amodal object segmentation is a challenging task that involves segmenting both visible and occluded parts of an object. In this paper, we propose a novel approach, called Coarse-to-Fine Segmentation (C2F-Seg), that addresses this problem by progressively modeling the amodal segmentation. C2F-Seg initially reduces the learning space from the pixel-level image space to the vector-quantized latent space. This enables us to better handle long-range dependencies and learn a coarse-grained amodal segment from visual features and visible segments. However, this latent space lacks detailed information about the object, which makes it difficult to provide a precise segmentation directly. To address this issue, we propose a convolution refine module to inject fine-grained information and provide a more precise amodal object segmentation based on visual features and coarse-predicted segmentation. To help the studies of amodal object segmentation, we create a synthetic amodal dataset, named as MOViD-Amodal (MOViD-A), which can be used for both image and video amodal object segmentation. We extensively evaluate our model on two benchmark datasets: KINS and COCO-A. Our empirical results demonstrate the superiority of C2F-Seg. Moreover, we exhibit the potential of our approach for video amodal object segmentation tasks on FISHBOWL and our proposed MOViD-A. Project page at: http://jianxgao.github.io/C2F-Seg.
翻译:非模态物体分割是一项具有挑战性的任务,涉及对物体可见部分和遮挡部分进行分割。本文提出了一种名为“由粗到精分割”(C2F-Seg)的新方法,通过渐进式建模非模态分割来解决该问题。C2F-Seg 首先将学习空间从像素级图像空间降维到向量量化潜在空间。这使我们能够更好地处理长程依赖关系,并从视觉特征和可见分割中学习粗粒度的非模态分割。然而,该潜在空间缺乏物体的细节信息,难以直接提供精确分割。为解决此问题,我们提出一个卷积精化模块,用于注入细粒度信息,并基于视觉特征和粗预测分割提供更精确的非模态物体分割。为促进非模态物体分割研究,我们创建了一个名为 MOViD-Amodal(MOViD-A)的合成非模态数据集,该数据集可用于图像和视频的非模态物体分割。我们在两个基准数据集(KINS 和 COCO-A)上全面评估了模型,实验结果证明了 C2F-Seg 的优越性。此外,我们在 FISHBOWL 和我们提出的 MOViD-A 上展示了该方法在视频非模态物体分割任务中的潜力。项目页面:http://jianxgao.github.io/C2F-Seg。