Self-supervised pre-training bears potential to generate expressive representations without human annotation. Most pre-training in Earth observation (EO) are based on ImageNet or medium-size, labeled remote sensing (RS) datasets. We share an unlabeled RS dataset SSL4EO-S12 (Self-Supervised Learning for Earth Observation - Sentinel-1/2) to assemble a large-scale, global, multimodal, and multi-seasonal corpus of satellite imagery from the ESA Sentinel-1 \& -2 satellite missions. For EO applications we demonstrate SSL4EO-S12 to succeed in self-supervised pre-training for a set of methods: MoCo-v2, DINO, MAE, and data2vec. Resulting models yield downstream performance close to, or surpassing accuracy measures of supervised learning. In addition, pre-training on SSL4EO-S12 excels compared to existing datasets. We make openly available the dataset, related source code, and pre-trained models at https://github.com/zhu-xlab/SSL4EO-S12.
翻译:自监督预训练能够在不依赖人工标注的情况下生成富有表达力的表示。当前地球观测领域的预训练主要基于ImageNet或中等规模的有标注遥感数据集。我们共享一个无标注遥感数据集SSL4EO-S12(面向地球观测的自监督学习——哨兵1/2号),该数据集由欧洲航天局哨兵1号和2号卫星任务获取的卫星影像组成,构建了大规模、全球覆盖、多模态且多季节的语料库。针对地球观测应用,我们证明了SSL4EO-S12在MoCo-v2、DINO、MAE和data2vec等一系列自监督预训练方法中均能取得良好效果。基于该数据集生成的模型在下游任务中的性能接近甚至超越监督学习的准确率指标。此外,在SSL4EO-S12上进行预训练的表现优于现有数据集。我们已在https://github.com/zhu-xlab/SSL4EO-S12 上开源提供该数据集、相关源代码及预训练模型。