Cross-modality medical image synthesis is a critical topic and has the potential to facilitate numerous applications in the medical imaging field. Despite recent successes in deep-learning-based generative models, most current medical image synthesis methods rely on generative adversarial networks and suffer from notorious mode collapse and unstable training. Moreover, the 2D backbone-driven approaches would easily result in volumetric inconsistency, while 3D backbones are challenging and impractical due to the tremendous memory cost and training difficulty. In this paper, we introduce a new paradigm for volumetric medical data synthesis by leveraging 2D backbones and present a diffusion-based framework, Make-A-Volume, for cross-modality 3D medical image synthesis. To learn the cross-modality slice-wise mapping, we employ a latent diffusion model and learn a low-dimensional latent space, resulting in high computational efficiency. To enable the 3D image synthesis and mitigate volumetric inconsistency, we further insert a series of volumetric layers in the 2D slice-mapping model and fine-tune them with paired 3D data. This paradigm extends the 2D image diffusion model to a volumetric version with a slightly increasing number of parameters and computation, offering a principled solution for generic cross-modality 3D medical image synthesis. We showcase the effectiveness of our Make-A-Volume framework on an in-house SWI-MRA brain MRI dataset and a public T1-T2 brain MRI dataset. Experimental results demonstrate that our framework achieves superior synthesis results with volumetric consistency.
翻译:跨模态医学图像合成是一个关键课题,有望为医学成像领域的诸多应用提供支持。尽管基于深度学习的生成模型近期取得了成功,但目前大多数医学图像合成方法仍依赖于生成对抗网络,且存在众所周知的模式坍塌和训练不稳定的问题。此外,基于2D主干的方法容易导致三维体积不一致性,而3D主干则因巨大的内存成本和训练难度而具有挑战性且不切实际。在本文中,我们介绍了一种利用2D主干进行体积医学数据合成的新范式,并提出了一种基于扩散的框架——Make-A-Volume,用于跨模态三维医学图像合成。为了学习跨模态切片级映射,我们采用潜在扩散模型并学习低维潜在空间,从而实现了高计算效率。为了实现三维图像合成并缓解体积不一致性,我们进一步在2D切片映射模型中插入一系列体积层,并使用配对的三维数据进行微调。该范式将2D图像扩散模型扩展为体积版本,且参数数量和计算量仅略有增加,为通用的跨模态三维医学图像合成提供了原则性解决方案。我们在内部SWI-MRA脑部核磁共振数据集和公开T1-T2脑部核磁共振数据集上展示了我们的Make-A-Volume框架的有效性。实验结果表明,我们的框架能够实现具有体积一致性的卓越合成效果。