Most change detection models based on vision transformers currently follow a "pretraining then fine-tuning" strategy. This involves initializing the model weights using large scale classification datasets, which can be either natural images or remote sensing images. However, fully tuning such a model requires significant time and resources. In this paper, we propose an efficient tuning approach that involves freezing the parameters of the pretrained image encoder and introducing additional training parameters. Through this approach, we have achieved competitive or even better results while maintaining extremely low resource consumption across six change detection benchmarks. For example, training time on LEVIR-CD, a change detection benchmark, is only half an hour with 9 GB memory usage, which could be very convenient for most researchers. Additionally, the decoupled tuning framework can be extended to any pretrained model for semantic change detection and multi temporal change detection as well. We hope that our proposed approach will serve as a part of foundational model to inspire more unified training approaches on change detection in the future.
翻译:当前基于视觉Transformer的变化检测模型大多遵循"预训练再微调"策略。该方法通过大规模分类数据集(包括自然图像或遥感图像)初始化模型权重。然而,对这类模型进行全参数微调需要耗费大量时间和资源。本文提出一种高效微调方法:冻结预训练图像编码器的参数,仅引入额外的训练参数。通过该方法,我们在六个变化检测基准上实现了极具竞争力的结果甚至更优性能,同时保持极低的资源消耗。例如,在变化检测基准LEVIR-CD上,训练时间仅需半小时,内存占用9 GB,这对大多数研究者而言极为便利。此外,这种解耦微调框架可扩展至任何预训练模型,用于语义变化检测和多时相变化检测。我们期望所提方法能作为基础模型的一部分,为未来变化检测领域更统一的训练方法提供启发。