Cross-modal medical image translation is an essential task for synthesizing missing modality data for clinical diagnosis. However, current learning-based techniques have limitations in capturing cross-modal and global features, restricting their suitability to specific pairs of modalities. This lack of versatility undermines their practical usefulness, particularly considering that the missing modality may vary for different cases. In this study, we present MedPrompt, a multi-task framework that efficiently translates different modalities. Specifically, we propose the Self-adaptive Prompt Block, which dynamically guides the translation network towards distinct modalities. Within this framework, we introduce the Prompt Extraction Block and the Prompt Fusion Block to efficiently encode the cross-modal prompt. To enhance the extraction of global features across diverse modalities, we incorporate the Transformer model. Extensive experimental results involving five datasets and four pairs of modalities demonstrate that our proposed model achieves state-of-the-art visual quality and exhibits excellent generalization capability.
翻译:跨模态医学图像翻译是临床诊断中合成缺失模态数据的关键任务。然而,现有基于学习的方法在捕获跨模态与全局特征方面存在局限性,导致其仅适用于特定模态对。这种通用性的缺失削弱了其实际应用价值——尤其考虑到不同病例中缺失的模态可能各异。本研究提出MedPrompt——一种高效翻译多种模态的多任务框架。具体而言,我们提出自适应提示块(Self-adaptive Prompt Block),可动态引导翻译网络针对不同模态进行适配。在该框架中,我们引入提示提取块(Prompt Extraction Block)与提示融合块(Prompt Fusion Block)以高效编码跨模态提示。为增强不同模态间全局特征的提取能力,我们集成了Transformer模型。涉及五个数据集与四组模态对的广泛实验结果表明,所提模型在视觉质量上达到最优水平,并展现出卓越的泛化能力。