The unstructured nature of clinical notes within electronic health records often conceals vital patient-related information, making it challenging to access or interpret. To uncover this hidden information, specialized Natural Language Processing (NLP) models are required. However, training these models necessitates large amounts of labeled data, a process that is both time-consuming and costly when relying solely on human experts for annotation. In this paper, we propose an approach that combines Large Language Models (LLMs) with human expertise to create an efficient method for generating ground truth labels for medical text annotation. By utilizing LLMs in conjunction with human annotators, we significantly reduce the human annotation burden, enabling the rapid creation of labeled datasets. We rigorously evaluate our method on a medical information extraction task, demonstrating that our approach not only substantially cuts down on human intervention but also maintains high accuracy. The results highlight the potential of using LLMs to improve the utilization of unstructured clinical data, allowing for the swift deployment of tailored NLP solutions in healthcare.
翻译:电子健康记录中临床笔记的非结构化特性往往掩盖了与患者相关的重要信息,使其难以访问或解读。为揭示这些隐藏信息,需借助专门的自然语言处理(NLP)模型。然而,训练此类模型需要大量标注数据,若仅依赖人类专家进行标注,这一过程既耗时又昂贵。本文提出一种将大语言模型(LLMs)与人类专业知识相结合的方法,旨在为医学文本标注生成真实标签的高效途径。通过将LLMs与人类标注员协同使用,我们显著降低了人工标注负担,从而能够快速构建标注数据集。我们在医学信息抽取任务上严格评估了该方法,结果表明我们的方法不仅大幅减少了人工干预,还保持了高精度。这些结果凸显了利用LLMs改善非结构化临床数据利用的潜力,从而助力医疗领域快速部署定制化NLP解决方案。