Radiology reports are detailed text descriptions of the content of medical scans. Each report describes the presence/absence and location of relevant clinical findings, commonly including comparison with prior exams of the same patient to describe how they evolved. Radiology reporting is a time-consuming process, and scan results are often subject to delays. One strategy to speed up reporting is to integrate automated reporting systems, however clinical deployment requires high accuracy and interpretability. Previous approaches to automated radiology reporting generally do not provide the prior study as input, precluding comparison which is required for clinical accuracy in some types of scans, and offer only unreliable methods of interpretability. Therefore, leveraging an existing visual input format of anatomical tokens, we introduce two novel aspects: (1) longitudinal representation learning -- we input the prior scan as an additional input, proposing a method to align, concatenate and fuse the current and prior visual information into a joint longitudinal representation which can be provided to the multimodal report generation model; (2) sentence-anatomy dropout -- a training strategy for controllability in which the report generator model is trained to predict only sentences from the original report which correspond to the subset of anatomical regions given as input. We show through in-depth experiments on the MIMIC-CXR dataset how the proposed approach achieves state-of-the-art results while enabling anatomy-wise controllable report generation.
翻译:放射学报告是对医学扫描内容提供详细文本描述的文件。每份报告描述相关临床发现的存在/缺失及位置,通常包括与同一患者既往检查结果的对比,以说明其演变过程。放射学报告撰写是一项耗时的工作,扫描结果常因报告延迟而受到影响。加速报告撰写的策略之一是集成自动化报告系统,然而临床部署要求高准确性和可解释性。先前的自动化放射学报告方法通常不提供既往研究作为输入,从而排除了某些类型扫描中临床准确性所需的对比分析,并且仅提供不可靠的可解释性方法。因此,利用已有的解剖学标记视觉输入格式,我们引入两项创新:(1)纵向表征学习——将既往扫描作为额外输入,提出一种对齐、拼接和融合当前与既往视觉信息的方法,形成可输入多模态报告生成模型的联合纵向表征;(2)句子-解剖学丢弃——一种用于可控性的训练策略,使报告生成模型仅预测原始报告中对应该输入解剖区域子集的句子。通过在MIMIC-CXR数据集上的深入实验,我们展示了所提方法在实现解剖学可控报告生成的同时,如何取得当前最优结果。