Image fusion aims to combine information from different source images to create a comprehensively representative image. Existing fusion methods are typically helpless in dealing with degradations in low-quality source images and non-interactive to multiple subjective and objective needs. To solve them, we introduce a novel approach that leverages semantic text guidance image fusion model for degradation-aware and interactive image fusion task, termed as Text-IF. It innovatively extends the classical image fusion to the text guided image fusion along with the ability to harmoniously address the degradation and interaction issues during fusion. Through the text semantic encoder and semantic interaction fusion decoder, Text-IF is accessible to the all-in-one infrared and visible image degradation-aware processing and the interactive flexible fusion outcomes. In this way, Text-IF achieves not only multi-modal image fusion, but also multi-modal information fusion. Extensive experiments prove that our proposed text guided image fusion strategy has obvious advantages over SOTA methods in the image fusion performance and degradation treatment. The code is available at https://github.com/XunpengYi/Text-IF.
翻译:图像融合旨在整合不同源图像中的信息,生成具有全面代表性的图像。现有融合方法通常难以处理低质量源图像中的退化问题,且无法适应多种主观与客观需求。为解决这些问题,我们提出了一种新颖方法——Text-IF,该方法利用语义文本引导的图像融合模型,实现退化感知与交互式图像融合任务。它创新地将传统图像融合扩展至文本引导的图像融合,并能够和谐解决融合过程中的退化与交互问题。通过文本语义编码器和语义交互融合解码器,Text-IF 实现了全能型红外与可见光图像退化感知处理,并产出交互式灵活融合结果。由此,Text-IF 不仅实现了多模态图像融合,还实现了多模态信息融合。大量实验证明,我们提出的文本引导图像融合策略在图像融合性能与退化处理方面,相较于最先进方法具有显著优势。代码已开源于 https://github.com/XunpengYi/Text-IF。