In this paper, we show that recent advances in self-supervised feature learning enable unsupervised object discovery and semantic segmentation with a performance that matches the state of the field on supervised semantic segmentation 10 years ago. We propose a methodology based on unsupervised saliency masks and self-supervised feature clustering to kickstart object discovery followed by training a semantic segmentation network on pseudo-labels to bootstrap the system on images with multiple objects. We present results on PASCAL VOC that go far beyond the current state of the art (50.0 mIoU), and we report for the first time results on MS COCO for the whole set of 81 classes: our method discovers 34 categories with more than $20\%$ IoU, while obtaining an average IoU of 19.6 for all 81 categories.
翻译:本文表明,自监督特征学习的最新进展实现了无监督对象发现与语义分割,其性能可与十年前监督语义分割领域的顶尖成果相媲美。我们提出一种方法论:先利用无监督显著性掩码和自监督特征聚类启动对象发现,再基于伪标签训练语义分割网络,从而在多对象图像上引导系统启动。我们在PASCAL VOC数据集上取得的成果远超当前最优水平(50.0 mIoU),并首次报告了在MS COCO全部81个类别上的结果:我们的方法发现了34个IoU超过20%的类别,所有81个类别的平均IoU达到19.6。