Understanding 3D medical image volumes is a critical task in the medical domain. However, existing 3D convolution and transformer-based methods have limited semantic understanding of an image volume and also need a large set of volumes for training. Recent advances in multi-modal large language models (MLLMs) provide a new and promising way to understand images with the help of text descriptions. However, most current MLLMs are designed for 2D natural images. To enhance the 3D medical image understanding with 2D MLLMs, we propose a novel pre-training framework called Med3DInsight, which marries existing 3D image encoders with 2D MLLMs and bridges them via a designed Plane-Slice-Aware Transformer (PSAT) module. Extensive experiments demonstrate our SOTA performance on two downstream segmentation and classification tasks, including three public datasets with CT and MRI modalities and comparison to more than ten baselines. Med3DInsight can be easily integrated into any current 3D medical image understanding network and improves its performance by a good margin.
翻译:理解三维医学图像体数据是医学领域的一项关键任务。然而,现有的基于三维卷积和Transformer的方法对图像体数据的语义理解有限,并且需要大量体数据进行训练。多模态大语言模型的最新进展提供了一种借助文本描述理解图像的全新且富有前景的方式。然而,当前大多数多模态大语言模型是专为二维自然图像设计的。为了利用二维多模态大语言模型增强三维医学图像理解,我们提出了一种名为Med3DInsight的新型预训练框架。该框架将现有三维图像编码器与二维多模态大语言模型相融合,并通过所设计的平面-切片感知Transformer模块连接二者。大量实验表明,我们在两个下游分割与分类任务中达到了最先进的性能,涉及包含CT和MRI模态的三个公开数据集,并与十多种基线方法进行了对比。Med3DInsight可轻松集成到现有任何三维医学图像理解网络中,并显著提升其性能。