The MEDIQA-CORR 2024 shared task aims to assess the ability of Large Language Models (LLMs) to identify and correct medical errors in clinical notes. In this study, we evaluate the capability of general LLMs, specifically GPT-3.5 and GPT-4, to identify and correct medical errors with multiple prompting strategies. Recognising the limitation of LLMs in generating accurate corrections only via prompting strategies, we propose incorporating error-span predictions from a smaller, fine-tuned model in two ways: 1) by presenting it as a hint in the prompt and 2) by framing it as multiple-choice questions from which the LLM can choose the best correction. We found that our proposed prompting strategies significantly improve the LLM's ability to generate corrections. Our best-performing solution with 8-shot + CoT + hints ranked sixth in the shared task leaderboard. Additionally, our comprehensive analyses show the impact of the location of the error sentence, the prompted role, and the position of the multiple-choice option on the accuracy of the LLM. This prompts further questions about the readiness of LLM to be implemented in real-world clinical settings.
翻译:MEDIQA-CORR 2024共享任务旨在评估大语言模型(LLMs)在临床记录中识别与纠正医学错误的能力。本研究通过多种提示策略,评估了通用大语言模型(特别是GPT-3.5与GPT-4)的医学错误识别与纠正能力。鉴于仅通过提示策略难以使大语言模型生成准确修正,我们提出将经微调的小型模型的错误跨度预测以两种方式整合:1)作为提示中的线索信息呈现;2)构建为多项选择题供大语言模型选择最佳修正。实验表明,我们提出的提示策略显著提升了大语言模型生成修正的能力。采用8样本+思维链+线索提示的最佳方案在共享任务排行榜中位列第六。此外,系统性分析揭示了错误句位置、提示角色设定及多项选择题选项排列对大语言模型准确率的影响。这进一步引发了对大语言模型在实际临床场景中应用成熟度的深入思考。