Accurately determining salient regions of an image is challenging when labeled data is scarce. DINO-based self-supervised approaches have recently leveraged meaningful image semantics captured by patch-wise features for locating foreground objects. Recent methods have also incorporated intuitive priors and demonstrated value in unsupervised methods for object partitioning. In this paper, we propose SEMPART, which jointly infers coarse and fine bi-partitions over an image's DINO-based semantic graph. Furthermore, SEMPART preserves fine boundary details using graph-driven regularization and successfully distills the coarse mask semantics into the fine mask. Our salient object detection and single object localization findings suggest that SEMPART produces high-quality masks rapidly without additional post-processing and benefits from co-optimizing the coarse and fine branches.
翻译:摘要:在标记数据稀缺的情况下,准确确定图像的显著区域具有挑战性。基于DINO的自监督方法近期利用补丁级特征捕获的有意义图像语义来定位前景物体。最新方法还融入了直观先验,并在无监督物体分割方法中展现出价值。本文提出SEMPART方法,该方法在基于DINO的图像语义图上联合推断粗粒度和细粒度二元分割。此外,SEMPART通过图驱动正则化保留精细边界细节,并成功将粗粒度掩码的语义信息蒸馏至细粒度掩码中。我们的显著性目标检测与单目标定位实验结果表明,SEMPART无需额外后处理即可快速生成高质量掩码,且能从粗粒度分支与细粒度分支的协同优化中获益。