In the field of medical Vision-Language Pre-training (VLP), significant efforts have been devoted to deriving text and image features from both clinical reports and associated medical images. However, most existing methods may have overlooked the opportunity in leveraging the inherent hierarchical structure of clinical reports, which are generally split into `findings' for descriptive content and `impressions' for conclusive observation. Instead of utilizing this rich, structured format, current medical VLP approaches often simplify the report into either a unified entity or fragmented tokens. In this work, we propose a novel clinical prior guided VLP framework named IMITATE to learn the structure information from medical reports with hierarchical vision-language alignment. The framework derives multi-level visual features from the chest X-ray (CXR) images and separately aligns these features with the descriptive and the conclusive text encoded in the hierarchical medical report. Furthermore, a new clinical-informed contrastive loss is introduced for cross-modal learning, which accounts for clinical prior knowledge in formulating sample correlations in contrastive learning. The proposed model, IMITATE, outperforms baseline VLP methods across six different datasets, spanning five medical imaging downstream tasks. Comprehensive experimental results highlight the advantages of integrating the hierarchical structure of medical reports for vision-language alignment.
翻译:在医学视觉-语言预训练(VLP)领域,大量研究致力于从临床报告及关联医学图像中提取文本与图像特征。然而,现有方法大多忽略了利用临床报告固有的分层结构——其通常分为用于描述性内容的“发现”和用于结论性观察的“印象”。当前医学VLP方法往往将报告简化为统一实体或碎片化词元,未能充分利用这一结构化格式。为此,我们提出一种名为IMITATE的新型临床先验引导VLP框架,通过学习医学报告中的结构信息实现分层视觉-语言对齐。该框架从胸部X光(CXR)图像中提取多层级视觉特征,并分别将其与分层医学报告中编码的描述性文本及结论性文本进行对齐。此外,我们引入一种新的临床信息对比损失用于跨模态学习,在对比学习中通过临床先验知识构建样本关联。所提模型IMITATE在涵盖五项医学影像下游任务的六个数据集上均优于基线VLP方法。全面实验结果凸显了融合医学报告分层结构对视觉-语言对齐的优越性。