Multimodal medical imaging plays a pivotal role in clinical diagnosis and research, as it combines information from various imaging modalities to provide a more comprehensive understanding of the underlying pathology. Recently, deep learning-based multimodal fusion techniques have emerged as powerful tools for improving medical image classification. This review offers a thorough analysis of the developments in deep learning-based multimodal fusion for medical classification tasks. We explore the complementary relationships among prevalent clinical modalities and outline three main fusion schemes for multimodal classification networks: input fusion, intermediate fusion (encompassing single-level fusion, hierarchical fusion, and attention-based fusion), and output fusion. By evaluating the performance of these fusion techniques, we provide insight into the suitability of different network architectures for various multimodal fusion scenarios and application domains. Furthermore, we delve into challenges related to network architecture selection, handling incomplete multimodal data management, and the potential limitations of multimodal fusion. Finally, we spotlight the promising future of Transformer-based multimodal fusion techniques and give recommendations for future research in this rapidly evolving field.
翻译:多模态医学成像在临床诊断与研究中扮演着关键角色,它通过整合来自不同成像模态的信息,从而更全面地理解潜在病理机制。近年来,基于深度学习的多模态融合技术已成为提升医学图像分类性能的有力工具。本综述深入分析了用于医学分类任务的深度学习多模态融合技术的发展历程。我们探讨了常见临床模态之间的互补关系,并概述了多模态分类网络的三种主要融合方案:输入融合、中间融合(包括单层级融合、层级融合和基于注意力的融合)以及输出融合。通过评估这些融合技术的性能,我们深入分析了不同网络架构在不同多模态融合场景及应用领域的适用性。此外,我们还探讨了与网络架构选择、不完整多模态数据处理以及多模态融合潜在局限性相关的挑战。最后,我们重点展望了基于Transformer的多模态融合技术的广阔前景,并针对这一快速发展领域提出了未来研究建议。