Multi-modality image fusion involves integrating complementary information from different modalities into a single image. Current methods primarily focus on enhancing image fusion with a single advanced task such as incorporating semantic or object-related information into the fusion process. This method creates challenges in achieving multiple objectives simultaneously. We introduce a target and semantic awareness joint-driven fusion network called TSJNet. TSJNet comprises fusion, detection, and segmentation subnetworks arranged in a series structure. It leverages object and semantically relevant information derived from dual high-level tasks to guide the fusion network. Additionally, We propose a local significant feature extraction module with a double parallel branch structure to fully capture the fine-grained features of cross-modal images and foster interaction among modalities, targets, and segmentation information. We conducted extensive experiments on four publicly available datasets (MSRS, M3FD, RoadScene, and LLVIP). The results demonstrate that TSJNet can generate visually pleasing fused results, achieving an average increase of 2.84% and 7.47% in object detection and segmentation mAP @0.5 and mIoU, respectively, compared to the state-of-the-art methods.
翻译:多模态图像融合旨在将来自不同模态的互补信息整合到单一图像中。当前方法主要侧重于通过单一高级任务(如将语义或目标相关信息融入融合过程)来增强图像融合效果,这种方法难以同时实现多个目标。我们提出了一种名为TSJNet的目标与语义联合感知驱动融合网络。TSJNet由融合、检测和分割子网络按串联结构组成,利用从双重高级任务中提取的目标与语义相关信息来引导融合网络。此外,我们提出了一种具有双并行分支结构的局部显著特征提取模块,以充分捕获跨模态图像的细粒度特征,并促进模态、目标和分割信息之间的交互。我们在四个公开数据集(MSRS、M3FD、RoadScene和LLVIP)上进行了广泛实验。结果表明,TSJNet能够生成视觉上令人满意的融合结果,与最先进方法相比,目标检测和分割分别在[email protected]和mIoU指标上平均提升2.84%和7.47%。