Co-salient Object Detection (CoSOD) endeavors to replicate the human visual system's capacity to recognize common and salient objects within a collection of images. Despite recent advancements in deep learning models, these models still rely on training with well-annotated CoSOD datasets. The exploration of training-free zero-shot CoSOD frameworks has been limited. In this paper, taking inspiration from the zero-shot transfer capabilities of foundational computer vision models, we introduce the first zero-shot CoSOD framework that harnesses these models without any training process. To achieve this, we introduce two novel components in our proposed framework: the group prompt generation (GPG) module and the co-saliency map generation (CMP) module. We evaluate the framework's performance on widely-used datasets and observe impressive results. Our approach surpasses existing unsupervised methods and even outperforms fully supervised methods developed before 2020, while remaining competitive with some fully supervised methods developed before 2022.
翻译:共显着目标检测旨在复制人类视觉系统从图像集合中识别常见且显着目标的能力。尽管深度学习模型近期取得了进展,但这些模型仍然依赖于使用标注完善的共显着目标检测数据集进行训练。对无需训练即可实现的零样本共显着目标检测框架的探索一直较为有限。本文受基础计算机视觉模型零样本迁移能力的启发,首次提出一种无需任何训练过程即可利用这些模型的零样本共显着目标检测框架。为实现这一目标,我们在所提出的框架中引入了两个新组件:组提示生成模块和共显着图生成模块。我们在广泛使用的数据集上评估了该框架的性能,并观察到了显著成果。我们的方法超越了现有的无监督方法,甚至优于2020年之前开发的某些全监督方法,同时与2022年之前开发的某些全监督方法相比仍具竞争力。