Generative Artificial Intelligence (AI) can be used to automatically generate medical reports based on transcripts of medical consultations. The aim is to reduce the administrative burden that healthcare professionals face. The accuracy of the generated reports needs to be established to ensure their correctness and usefulness. There are several metrics for measuring the accuracy of AI generated reports, but little work has been done towards the application of these metrics in medical reporting. A comparative experimentation of 10 accuracy metrics has been performed on AI generated medical reports against their corresponding General Practitioner's (GP) medical reports concerning Otitis consultations. The number of missing, incorrect, and additional statements of the generated reports have been correlated with the metric scores. In addition, we introduce and define a Composite Accuracy Score which produces a single score for comparing the metrics within the field of automated medical reporting. Findings show that based on the correlation study and the Composite Accuracy Score, the ROUGE-L and Word Mover's Distance metrics are the preferred metrics, which is not in line with previous work. These findings help determine the accuracy of an AI generated medical report, which aids the development of systems that generate medical reports for GPs to reduce the administrative burden.
翻译:生成式人工智能可用于根据医疗咨询记录自动生成医学报告,其目标是减轻医疗专业人员面临的行政负担。生成的报告需要确定其准确性,以确保正确性和实用性。目前存在多种用于衡量AI生成报告准确率的指标,但针对这些指标在医疗报告中的应用研究尚不充分。本研究对AI生成的医疗报告进行了10项准确率指标的对比实验,并与相应全科医生针对耳炎咨询撰写的医疗报告进行对照。我们将生成报告中缺失、错误和额外陈述的数量与指标得分进行了相关性分析。此外,我们引入并定义了一个复合准确率得分,该得分可生成单一数值用于比较自动化医疗报告领域的各项指标。研究结果表明,基于相关性分析和复合准确率得分,ROUGE-L和词移动距离指标是更优选的指标,这与先前研究结论不一致。这些发现有助于确定AI生成医疗报告的准确率,从而支持开发为全科医生生成医疗报告以减轻行政负担的系统。