Background: Biomedical entity normalization is critical to biomedical research because the richness of free-text clinical data, such as progress notes, can often be fully leveraged only after translating words and phrases into structured and coded representations suitable for analysis. Large Language Models (LLMs), in turn, have shown great potential and high performance in a variety of natural language processing (NLP) tasks, but their application for normalization remains understudied. Methods: We applied both proprietary and open-source LLMs in combination with several rule-based normalization systems commonly used in biomedical research. We used a two-step LLM integration approach, (1) using an LLM to generate alternative phrasings of a source utterance, and (2) to prune candidate UMLS concepts, using a variety of prompting methods. We measure results by $F_{\beta}$, where we favor recall over precision, and F1. Results: We evaluated a total of 5,523 concept terms and text contexts from a publicly available dataset of human-annotated biomedical abstracts. Incorporating GPT-3.5-turbo increased overall $F_{\beta}$ and F1 in normalization systems +9.5 and +7.3 (MetaMapLite), +13.9 and +10.9 (QuickUMLS), and +10.5 and +10.3 (BM25), while the open-source Vicuna model achieved +10.8 and +12.2 (MetaMapLite), +14.7 and +15 (QuickUMLS), and +15.6 and +18.7 (BM25). Conclusions: Existing general-purpose LLMs, both propriety and open-source, can be leveraged at scale to greatly improve normalization performance using existing tools, with no fine-tuning.
翻译:背景:生物医学实体规范化对生物医学研究至关重要,因为自由文本临床数据(如病程记录)的丰富性通常只有在将词语和短语转化为适合分析的结构化编码表示后才能被充分利用。大型语言模型(LLMs)已在多种自然语言处理(NLP)任务中展现出巨大潜力与高性能,但其在规范化任务中的应用仍待深入探索。方法:我们将专有与开源LLMs与生物医学研究中常用的多种基于规则的规范化系统相结合,采用两步式LLM集成策略:(1)使用LLM生成源表述的替代表达形式,(2)通过多种提示方法对候选UMLS概念进行筛选。我们采用$F_{\beta}$(侧重召回率而非精确率)和F1分数评估结果。结果:我们在公开的人类标注生物医学摘要数据集中评估了总计5,523个概念术语及文本上下文。集成GPT-3.5-turbo使规范化系统的整体$F_{\beta}$和F1分别提升:+9.5与+7.3(MetaMapLite)、+13.9与+10.9(QuickUMLS)、+10.5与+10.3(BM25);而开源Vicuna模型则实现:+10.8与+12.2(MetaMapLite)、+14.7与+15(QuickUMLS)、+15.6与+18.7(BM25)。结论:现有通用型LLMs(包括专有与开源模型)无需微调即可规模化集成至现有工具,从而显著提升规范化性能。