Foundation models have excelled in various tasks but are often evaluated on general benchmarks. The adaptation of these models for specific domains, such as remote sensing imagery, remains an underexplored area. In remote sensing, precise building instance segmentation is vital for applications like urban planning. While Convolutional Neural Networks (CNNs) perform well, their generalization can be limited. For this aim, we present a novel approach to adapt foundation models to address existing models' generalization dropback. Among several models, our focus centers on the Segment Anything Model (SAM), a potent foundation model renowned for its prowess in class-agnostic image segmentation capabilities. We start by identifying the limitations of SAM, revealing its suboptimal performance when applied to remote sensing imagery. Moreover, SAM does not offer recognition abilities and thus fails to classify and tag localized objects. To address these limitations, we introduce different prompting strategies, including integrating a pre-trained CNN as a prompt generator. This novel approach augments SAM with recognition abilities, a first of its kind. We evaluated our method on three remote sensing datasets, including the WHU Buildings dataset, the Massachusetts Buildings dataset, and the AICrowd Mapping Challenge. For out-of-distribution performance on the WHU dataset, we achieve a 5.47% increase in IoU and a 4.81% improvement in F1-score. For in-distribution performance on the WHU dataset, we observe a 2.72% and 1.58% increase in True-Positive-IoU and True-Positive-F1 score, respectively. We intend to release our code repository, hoping to inspire further exploration of foundation models for domain-specific tasks within the remote sensing community.
翻译:基础模型在各种任务中表现出色,但通常仅在通用基准上进行评估。这些模型在特定领域(如遥感影像)中的适应性仍是一个未充分探索的领域。在遥感领域,精确的建筑物实例分割对城市规æ规划等应用至关重要。尽管卷积神经网络(CNN)表现良好,但其泛化能力可能有限。为此,我们提出了一种新颖的方法来调整基础模型,以解决现有模型的泛化退化问题。在众多模型中,我们重点关注分割任意模型(SAM),这是一种强大的基础模型,以其在类别无关图像分割方面的能力而闻名。我们首先识别了SAM的局限性,揭示其在遥感影像应用中的次优表现。此外,SAM不具备识别能力,因此无法对定位目标进行分类和标注。为解决这些问题,我们引入了多种提示策略,包括集成预训练的CNN作为提示生成器。这一新颖方法首次为SAM赋予了识别能力。我们在三个遥感数据集上评估了方法,包括WHU建筑物数据集、马萨诸塞州建筑物数据集和AICrowd地图挑战赛。在WHU数据集上的分布外性能中,IoU提高了5.47%,F1分数提升了4.81%。在WHU数据集上的分布内性能中,True-Positive-IoU和True-Positive-F1分数分别提升了2.72%和1.58%。我们计划发布代码仓库,希望激发遥感社区进一步探索面向特定领域任务的基础模型应用。