Both medical care and observational studies in oncology require a thorough understanding of a patient's disease progression and treatment history, often elaborately documented in clinical notes. Despite their vital role, no current oncology information representation and annotation schema fully encapsulates the diversity of information recorded within these notes. Although large language models (LLMs) have recently exhibited impressive performance on various medical natural language processing tasks, due to the current lack of comprehensively annotated oncology datasets, an extensive evaluation of LLMs in extracting and reasoning with the complex rhetoric in oncology notes remains understudied. We developed a detailed schema for annotating textual oncology information, encompassing patient characteristics, tumor characteristics, tests, treatments, and temporality. Using a corpus of 40 de-identified breast and pancreatic cancer progress notes at University of California, San Francisco, we applied this schema to assess the zero-shot abilities of three recent LLMs (GPT-4, GPT-3.5-turbo, and FLAN-UL2) to extract detailed oncological history from two narrative sections of clinical progress notes. Our team annotated 9028 entities, 9986 modifiers, and 5312 relationships. The GPT-4 model exhibited overall best performance, with an average BLEU score of 0.73, an average ROUGE score of 0.72, an exact-match F1-score of 0.51, and an average accuracy of 68% on complex tasks (expert manual evaluation on subset). Notably, it was proficient in tumor characteristic and medication extraction, and demonstrated superior performance in relational inference like adverse event detection. However, further improvements are needed before using it to reliably extract important facts from cancer progress notes needed for clinical research, complex population management, and documenting quality patient care.
翻译:肿瘤学的医疗护理和观察性研究均需全面理解患者疾病进展和治疗史,这些信息通常详细记录于临床病程中。尽管临床病程至关重要,但当前尚无肿瘤学信息表示与标注模式能完全涵盖其中记录的信息多样性。尽管大语言模型近期在多种医学自然语言处理任务中展现出卓越性能,但由于缺乏全面标注的肿瘤学数据集,目前对LLMs在提取和推理肿瘤学病程中复杂修辞结构的系统性评估仍不充分。我们开发了涵盖患者特征、肿瘤特征、检测、治疗及时序性的肿瘤学文本信息详细标注模式。基于加州大学旧金山分校40份去标识化的乳腺癌和胰腺癌病程记录语料库,我们应用该模式评估了三种最新LLMs(GPT-4、GPT-3.5-turbo和FLAN-UL2)从临床病程两个叙述段落中提取详细肿瘤病史的零样本能力。团队共标注了9028个实体、9986个修饰语和5312个关系。GPT-4模型展现了整体最佳性能,其平均BLEU得分为0.73,平均ROUGE得分为0.72,精确匹配F1得分为0.51,复杂任务平均准确率为68%(专家人工评估子集)。值得注意的是,该模型在肿瘤特征和药物提取方面表现优异,并在不良事件检测等关系推理中展现出卓越性能。但若要将其可靠地应用于临床研究、复杂人群管理及优质患者护理记录所需的关键事实提取,仍需进一步改进。