Although supervised machine learning is popular for information extraction from clinical notes, creating large annotated datasets requires extensive domain expertise and is time-consuming. Meanwhile, large language models (LLMs) have demonstrated promising transfer learning capability. In this study, we explored whether recent LLMs can reduce the need for large-scale data annotations. We curated a manually-labeled dataset of 769 breast cancer pathology reports, labeled with 13 categories, to compare zero-shot classification capability of the GPT-4 model and the GPT-3.5 model with supervised classification performance of three model architectures: random forests classifier, long short-term memory networks with attention (LSTM-Att), and the UCSF-BERT model. Across all 13 tasks, the GPT-4 model performed either significantly better than or as well as the best supervised model, the LSTM-Att model (average macro F1 score of 0.83 vs. 0.75). On tasks with high imbalance between labels, the differences were more prominent. Frequent sources of GPT-4 errors included inferences from multiple samples and complex task design. On complex tasks where large annotated datasets cannot be easily collected, LLMs can reduce the burden of large-scale data labeling. However, if the use of LLMs is prohibitive, the use of simpler supervised models with large annotated datasets can provide comparable results. LLMs demonstrated the potential to speed up the execution of clinical NLP studies by reducing the need for curating large annotated datasets. This may result in an increase in the utilization of NLP-based variables and outcomes in observational clinical studies.
翻译:尽管监督式机器学习在从临床记录中提取信息方面广受欢迎,但创建大规模标注数据集需要大量领域专业知识且耗时费力。与此同时,大型语言模型(LLMs)已展现出显著的迁移学习能力。本研究探索了当前LLMs能否减少对大规模数据标注的需求。我们整理了一个由769份乳腺癌病理报告组成的手动标注数据集,包含13个分类标签,以比较GPT-4模型和GPT-3.5模型的零样本分类能力与三种监督式模型架构的分类性能:随机森林分类器、基于注意力机制的长短期记忆网络(LSTM-Att)和UCSF-BERT模型。在所有13个任务中,GPT-4模型的性能要么显著优于最优监督式模型(LSTM-Att模型),要么与之相当(平均宏F1分数分别为0.83和0.75)。在标签高度不平衡的任务中,差异更为显著。GPT-4出错的常见原因包括基于多个样本的推断和复杂的任务设计。在难以轻松收集大规模标注数据集的复杂任务中,LLMs可降低大规模数据标注的负担。但如果LLMs的使用受限,采用具备大规模标注数据集的简单监督式模型也能获得可比结果。LLMs通过减少对大规模标注数据集的需求,展现了加快临床自然语言处理(NLP)研究执行的潜力。这可能促使观察性临床研究中基于NLP的变量和结局指标应用得到提升。