Background: Recent advancements in large language models (LLMs) offer potential benefits in healthcare, particularly in processing extensive patient records. However, existing benchmarks do not fully assess LLMs' capability in handling real-world, lengthy clinical data. Methods: We present the LongHealth benchmark, comprising 20 detailed fictional patient cases across various diseases, with each case containing 5,090 to 6,754 words. The benchmark challenges LLMs with 400 multiple-choice questions in three categories: information extraction, negation, and sorting, challenging LLMs to extract and interpret information from large clinical documents. Results: We evaluated nine open-source LLMs with a minimum of 16,000 tokens and also included OpenAI's proprietary and cost-efficient GPT-3.5 Turbo for comparison. The highest accuracy was observed for Mixtral-8x7B-Instruct-v0.1, particularly in tasks focused on information retrieval from single and multiple patient documents. However, all models struggled significantly in tasks requiring the identification of missing information, highlighting a critical area for improvement in clinical data interpretation. Conclusion: While LLMs show considerable potential for processing long clinical documents, their current accuracy levels are insufficient for reliable clinical use, especially in scenarios requiring the identification of missing information. The LongHealth benchmark provides a more realistic assessment of LLMs in a healthcare setting and highlights the need for further model refinement for safe and effective clinical application. We make the benchmark and evaluation code publicly available.
翻译:背景:大语言模型的最新进展为医疗领域带来潜在效益,尤其在处理海量患者记录方面。然而现有基准未能充分评估大语言模型处理真实长篇幅临床数据的能力。方法:我们提出LongHealth基准,包含20个跨多种疾病的虚构详细患者病例,每个病例含5,090至6,754个词。该基准通过400道涉及信息抽取、否定判断与排序三类任务的多项选择题,检验大语言模型从大型临床文档中提取与解读信息的能力。结果:我们评估了九个至少支持16,000个词元(tokens)的开源大语言模型,并将OpenAI专有且具成本效益的GPT-3.5 Turbo纳入对比。Mixtral-8x7B-Instruct-v0.1在单文档与多文档信息检索任务中准确率最高。然而,所有模型在识别缺失信息任务中均表现显著不足,凸显临床数据解读中亟待改进的关键领域。结论:尽管大语言模型在长临床文档处理上展现潜力,但其当前准确率尚不足以支撑可靠的临床实践,尤其在需识别缺失信息的场景中。LongHealth基准为医疗环境下的大语言模型提供了更真实的评估,并揭示了实现安全有效临床应用仍需进一步优化的方向。我们将基准与评估代码公开发布。