In recent years, the foundation models have swept the computer vision field and facilitated the development of various tasks within different modalities. However, it remains an open question on how to design an infrared foundation model. In this paper, we propose InfMAE, a foundation model in infrared modality. We release an infrared dataset, called Inf30 to address the problem of lacking large-scale data for self-supervised learning in the infrared vision community. Besides, we design an information-aware masking strategy, which is suitable for infrared images. This masking strategy allows for a greater emphasis on the regions with richer information in infrared images during the self-supervised learning process, which is conducive to learning the generalized representation. In addition, we adopt a multi-scale encoder to enhance the performance of the pre-trained encoders in downstream tasks. Finally, based on the fact that infrared images do not have a lot of details and texture information, we design an infrared decoder module, which further improves the performance of downstream tasks. Extensive experiments show that our proposed method InfMAE outperforms other supervised methods and self-supervised learning methods in three downstream tasks. Our code will be made public at https://github.com/liufangcen/InfMAE.
翻译:近年来,基础模型席卷了计算机视觉领域,并推动了不同模态下各类任务的发展。然而,如何设计红外基础模型仍是一个悬而未决的问题。本文提出InfMAE,一种红外模态的基础模型。我们发布了一个名为Inf30的红外数据集,以解决红外视觉社区中缺乏大规模自监督学习数据的问题。此外,我们设计了一种适用于红外图像的信息感知掩码策略。该掩码策略使得自监督学习过程中能够更关注红外图像中信息更丰富的区域,有助于学习通用表征。同时,我们采用多尺度编码器来增强预训练编码器在下游任务中的性能。最后,基于红外图像缺乏大量细节和纹理信息这一事实,我们设计了一个红外解码器模块,进一步提升了下游任务的性能。大量实验表明,我们提出的InfMAE方法在三个下游任务中优于其他有监督方法和自监督学习方法。我们的代码将在https://github.com/liufangcen/InfMAE上公开。