The emergence of multi-modal deep learning models has made significant impacts on clinical applications in the last decade. However, the majority of models are limited to single-tasking, without considering disease diagnosis is indeed a multi-task procedure. Here, we demonstrate a unified transformer model specifically designed for multi-modal clinical tasks by incorporating customized instruction tuning. We first compose a multi-task training dataset comprising 13.4 million instruction and ground-truth pairs (with approximately one million radiographs) for the customized tuning, involving both image- and pixel-level tasks. Thus, we can unify the various vision-intensive tasks in a single training framework with homogeneous model inputs and outputs to increase clinical interpretability in one reading. Finally, we demonstrate the overall superior performance of our model compared to prior arts on various chest X-ray benchmarks across multi-tasks in both direct inference and finetuning settings. Three radiologists further evaluate the generated reports against the recorded ones, which also exhibit the enhanced explainability of our multi-task model.
翻译:多模态深度学习模型的出现在过去十年中对临床应用产生了显著影响。然而,大多数模型局限于单任务处理,未考虑疾病诊断本质上是一个多任务过程。本文提出了一种专为多模态临床任务设计的统一Transformer模型,通过融入定制化指令微调实现。我们首先构建了一个包含1340万条指令与真实标签对(约100万张X光片)的多任务训练数据集用于定制化微调,涵盖图像级和像素级任务。由此,我们可在单一训练框架中统一各类视觉密集型任务,通过同质化的模型输入与输出增强单次读片中的临床可解释性。最后,我们在多任务直接推理与微调场景下的各类胸部X光基准测试中,展示了模型相较于现有技术的总体优越性能。三位放射科医师进一步将模型生成的报告与原始记录进行对比评估,结果证实了本多任务模型增强的可解释性。