Medical vision-and-language pre-training (Med-VLP) has shown promising improvements on many downstream medical tasks owing to its applicability to extracting generic representations from medical images and texts. Practically, there exist two typical types, \textit{i.e.}, the fusion-encoder type and the dual-encoder type, depending on whether a heavy fusion module is used. The former is superior at multi-modal tasks owing to the sufficient interaction between modalities; the latter is good at uni-modal and cross-modal tasks due to the single-modality encoding ability. To take advantage of these two types, we propose an effective yet straightforward scheme named PTUnifier to unify the two types. We first unify the input format by introducing visual and textual prompts, which serve as a feature bank that stores the most representative images/texts. By doing so, a single model could serve as a \textit{foundation model} that processes various tasks adopting different input formats (\textit{i.e.}, image-only, text-only, and image-text-pair). Furthermore, we construct a prompt pool (instead of static ones) to improve diversity and scalability. Experimental results show that our approach achieves state-of-the-art results on a broad range of tasks, spanning uni-modal tasks (\textit{i.e.}, image/text classification and text summarization), cross-modal tasks (\textit{i.e.}, image-to-text generation and image-text/text-image retrieval), and multi-modal tasks (\textit{i.e.}, visual question answering), demonstrating the effectiveness of our approach. Note that the adoption of prompts is orthogonal to most existing Med-VLP approaches and could be a beneficial and complementary extension to these approaches.
翻译:医学视觉与语言预训练(Med-VLP)因其在从医学图像和文本中提取通用表示方面的适用性,已在诸多下游医学任务中展现出显著的性能提升。实践中存在两种典型类型,即融合编码器类型和双编码器类型,其区别在于是否使用重型融合模块。前者因模态间的充分交互而在多模态任务中表现优越;后者则凭借单模态编码能力,擅长处理单模态和跨模态任务。为综合两类方法的优势,我们提出一种名为PTUnifier的有效且简洁的统一方案。首先通过引入视觉和文本提示来统一输入格式,这些提示作为存储最具代表性图像/文本的特征库。由此,单一模型可作为处理多种任务(采用不同输入格式:仅图像、仅文本、图像-文本对)的《基础模型》。此外,我们构建了提示池(而非静态提示)以提升多样性与可扩展性。实验结果表明,我们的方法在涵盖单模态任务(即图像/文本分类与文本摘要)、跨模态任务(即图像到文本生成与图像-文本/文本-图像检索)以及多模态任务(即视觉问答)的广泛任务上均取得了最先进成果,验证了方法的有效性。值得注意的是,提示的采用与大多数现有Med-VLP方法正交,可作为这些方法的有效补充与延伸。