Visual foundation models have achieved remarkable results in zero-shot image classification and segmentation, but zero-shot change detection remains an open problem. In this paper, we propose the segment any change models (AnyChange), a new type of change detection model that supports zero-shot prediction and generalization on unseen change types and data distributions. AnyChange is built on the segment anything model (SAM) via our training-free adaptation method, bitemporal latent matching. By revealing and exploiting intra-image and inter-image semantic similarities in SAM's latent space, bitemporal latent matching endows SAM with zero-shot change detection capabilities in a training-free way. We also propose a point query mechanism to enable AnyChange's zero-shot object-centric change detection capability. We perform extensive experiments to confirm the effectiveness of AnyChange for zero-shot change detection. AnyChange sets a new record on the SECOND benchmark for unsupervised change detection, exceeding the previous SOTA by up to 4.4% F$_1$ score, and achieving comparable accuracy with negligible manual annotations (1 pixel per image) for supervised change detection.
翻译:视觉基础模型在零样本图像分类和分割中取得了显著成果,但零样本变化检测仍是一个开放性问题。本文提出"分割任何变化"模型(AnyChange),这是一种新型变化检测模型,支持对未见变化类型和数据分布的零样本预测与泛化。AnyChange基于分割任何模型(SAM),通过我们提出的免训练适应方法——双时相潜空间匹配构建而成。通过揭示并利用SAM潜空间中图像内与图像间的语义相似性,双时相潜空间匹配以免训练方式赋予SAM零样本变化检测能力。我们还提出点查询机制,使AnyChange具备零样本目标中心变化检测能力。通过大量实验验证了AnyChange在零样本变化检测中的有效性。在SECOND基准无监督变化检测任务中,AnyChange创下新纪录,F$_1$分数较先前最优结果提升最高达4.4%,并在监督变化检测中通过极少量人工标注(每张图像1像素)即可达到相当的精度。