With the availability of large-scale, comprehensive, and general-purpose vision-language (VL) datasets such as MSCOCO, vision-language pre-training (VLP) has become an active area of research and proven to be effective for various VL tasks such as visual-question answering. However, studies on VLP in the medical domain have so far been scanty. To provide a comprehensive perspective on VLP for medical VL tasks, we conduct a thorough experimental analysis to study key factors that may affect the performance of VLP with a unified vision-language Transformer. To allow making sound and quick pre-training decisions, we propose RadioGraphy Captions (RGC), a high-quality, multi-modality radiographic dataset containing 18,434 image-caption pairs collected from an open-access online database MedPix. RGC can be used as a pre-training dataset or a new benchmark for medical report generation and medical image-text retrieval. By utilizing RGC and other available datasets for pre-training, we develop several key insights that can guide future medical VLP research and new strong baselines for various medical VL tasks.
翻译:随着MSCOCO等大规模、综合性通用视觉-语言(VL)数据集的出现,视觉-语言预训练(VLP)已成为活跃的研究领域,并在视觉问答等多种VL任务中展现出显著效果。然而,医学领域的VLP研究至今仍然匮乏。为全面探究VLP在医学VL任务中的表现,我们通过统一视觉-语言Transformer进行深入实验分析,系统研究了影响VLP性能的关键因素。为支持快速合理的预训练决策,我们提出高质量多模态放射学数据集RadioGraphy Captions(RGC),该数据集包含从开放获取在线数据库MedPix收集的18,434对图像-文本描述。RGC可作为预训练数据集,也可作为医学报告生成与医学图像-文本检索的新基准。通过利用RGC及其他可用数据集进行预训练,我们总结出若干关键洞见,为未来医学VLP研究提供指导,并为各类医学VL任务建立新的强基线方法。