Radiology reports are an instrumental part of modern medicine, informing key clinical decisions such as diagnosis and treatment. The worldwide shortage of radiologists, however, restricts access to expert care and imposes heavy workloads, contributing to avoidable errors and delays in report delivery. While recent progress in automated report generation with vision-language models offer clear potential in ameliorating the situation, the path to real-world adoption has been stymied by the challenge of evaluating the clinical quality of AI-generated reports. In this study, we build a state-of-the-art report generation system for chest radiographs, \textit{Flamingo-CXR}, by fine-tuning a well-known vision-language foundation model on radiology data. To evaluate the quality of the AI-generated reports, a group of 16 certified radiologists provide detailed evaluations of AI-generated and human written reports for chest X-rays from an intensive care setting in the United States and an inpatient setting in India. At least one radiologist (out of two per case) preferred the AI report to the ground truth report in over 60$\%$ of cases for both datasets. Amongst the subset of AI-generated reports that contain errors, the most frequently cited reasons were related to the location and finding, whereas for human written reports, most mistakes were related to severity and finding. This disparity suggested potential complementarity between our AI system and human experts, prompting us to develop an assistive scenario in which \textit{Flamingo-CXR} generates a first-draft report, which is subsequently revised by a clinician. This is the first demonstration of clinician-AI collaboration for report writing, and the resultant reports are assessed to be equivalent or preferred by at least one radiologist to reports written by experts alone in 80$\%$ of in-patient cases and 60$\%$ of intensive care cases.
翻译:放射学报告是现代医学的重要组成部分,为诊断和治疗等关键临床决策提供依据。然而,全球放射科医生短缺限制了患者获得专家诊疗的机会,并导致放射科医生工作负荷过重,造成可预防的错误和报告延迟。尽管近期利用视觉-语言模型自动生成报告的进展为改善这一状况提供了明显潜力,但通往实际应用的道路仍受阻于评估AI生成报告临床质量的挑战。本研究通过微调一个知名的视觉-语言基础模型于放射学数据,构建了用于胸片报告生成的最先进系统Flamingo-CXR。为评估AI生成报告的质量,一组由16名认证放射科医生组成的评估团队对来自美国重症监护病房和印度住院病房的胸部X光片,详细评估了AI生成报告与人工撰写报告的质量。在两个数据集中,超过60%的病例(每位病例由两名放射科医生评估)至少有1名评估者认为AI报告优于真实报告。在包含错误的AI生成子集中,最常见的错误原因是与位置和发现相关;而在人工书写报告中,多数错误与严重程度和发现相关。这种差异表明AI系统与人类专家之间存在潜在互补性,促使我们开发一种辅助模式:由Flamingo-CXR生成初稿报告,再由临床医生进行修订。这是首次展示临床医生与AI协同撰写报告,评估结果显示:在80%的住院病例和60%的重症监护病例中,至少有1名放射科医生认为协同生成的报告与仅由专家独立撰写的报告质量相当或更优。