Open-vocabulary change detection aims to identify semantic changes in bi-temporal remote sensing images without predefined categories. Recent methods combine foundation models such as SAM, DINO and CLIP, but typically process each timestamp independently or interact only at the final comparison stage. Such paradigms suffer from insufficient temporal coupling during semantic reasoning, which limits their ability to distinguish genuine semantic changes from non-semantic appearance discrepancies. In addition, patch-dominant inference on high-resolution images often weakens global semantic continuity and produces fragmented change regions. To address these issues, we propose MemOVCD, a training-free open-vocabulary change detection framework based on cross-temporal memory reasoning and global-local adaptive rectification. Specifically, we reformulate bi-temporal change detection as a two-frame tracking problem and introduce weighted bidirectional propagation to aggregate semantic evidence from both temporal directions. To stabilize memory propagation across large temporal gaps, we construct histogram-aligned transition frames to smooth abrupt appearance changes. Moreover, a global-local adaptive rectification strategy adaptively fuses local and global-view predictions, improving spatial consistency while preserving fine-grained details. Experiments on five benchmarks demonstrate that MemOVCD achieves favorable performance on two change detection tasks, validating its effectiveness and generalization under diverse open-vocabulary settings.
翻译:开放词汇变化检测旨在无预定义类别条件下识别双时相遥感图像中的语义变化。现有方法通常结合SAM、DINO、CLIP等基础模型,但往往独立处理各时间节点或仅在最终对比阶段进行交互。此类范式在语义推理过程中存在时域耦合不足的问题,限制了其区分真正语义变化与非语义外观差异的能力。此外,针对高分辨率图像的分块主导推理方式常会削弱全局语义连续性,产生碎片化变化区域。为解决上述问题,我们提出MemOVCD——一种基于跨时域记忆推理与全局-局部自适应校正的无训练开放词汇变化检测框架。具体而言,我们将双时相变化检测重新建模为双帧跟踪问题,并引入加权双向传播机制以汇聚两个时相方向的语义证据。为稳定跨大时间间隔的记忆传播,我们构建直方图对齐过渡帧以平滑突发外观变化。进一步地,全局-局部自适应校正策略可自适应融合局部与全局视图预测结果,在保持细粒度细节的同时提升空间一致性。在五个基准数据集上的实验表明,MemOVCD在两项变化检测任务中均取得优异性能,验证了其在多样化开放词汇场景下的有效性与泛化能力。