Large language models (LLMs) have shown remarkable capabilities in Natural Language Processing (NLP), especially in domains where labeled data is scarce or expensive, such as clinical domain. However, to unlock the clinical knowledge hidden in these LLMs, we need to design effective prompts that can guide them to perform specific clinical NLP tasks without any task-specific training data. This is known as in-context learning, which is an art and science that requires understanding the strengths and weaknesses of different LLMs and prompt engineering approaches. In this paper, we present a comprehensive and systematic experimental study on prompt engineering for five clinical NLP tasks: Clinical Sense Disambiguation, Biomedical Evidence Extraction, Coreference Resolution, Medication Status Extraction, and Medication Attribute Extraction. We assessed the prompts proposed in recent literature, including simple prefix, simple cloze, chain of thought, and anticipatory prompts, and introduced two new types of prompts, namely heuristic prompting and ensemble prompting. We evaluated the performance of these prompts on three state-of-the-art LLMs: GPT-3.5, BARD, and LLAMA2. We also contrasted zero-shot prompting with few-shot prompting, and provide novel insights and guidelines for prompt engineering for LLMs in clinical NLP. To the best of our knowledge, this is one of the first works on the empirical evaluation of different prompt engineering approaches for clinical NLP in this era of generative AI, and we hope that it will inspire and inform future research in this area.
翻译:大语言模型在自然语言处理中展现出卓越能力,尤其在标注数据稀缺或昂贵的临床领域。然而,要释放这些大语言模型中隐藏的临床知识,需要设计有效的提示,引导它们在无任务特定训练数据的情况下执行具体临床自然语言处理任务。这被称为上下文学习,是一门需要理解不同大语言模型和提示工程方法优缺点的艺术与科学。本文对五种临床自然语言处理任务(临床词义消歧、生物医学证据提取、指代消解、用药状态提取和用药属性提取)的提示工程进行了全面且系统的实验研究。我们评估了近期文献中提出的提示方法,包括简单前缀、简单完形填空、思维链和预期提示,并引入了两种新型提示:启发式提示和集成提示。我们在三种最先进的大语言模型(GPT-3.5、BARD和LLAMA2)上评估了这些提示的性能,同时对比了零样本提示与少样本提示,为临床自然语言处理中大语言模型的提示工程提供了新颖见解和指导原则。据我们所知,这是当下生成式人工智能时代最早针对临床自然语言处理不同提示工程方法进行实证评估的研究之一,期望能够启发性指导该领域的未来研究方向。