Clinical text is rich in information, with mentions of treatment, medication and anatomy among many other clinical terms. Multiple terms can refer to the same core concepts which can be referred as a clinical entity. Ontologies like the Unified Medical Language System (UMLS) are developed and maintained to store millions of clinical entities including the definitions, relations and other corresponding information. These ontologies are used for standardization of clinical text by normalizing varying surface forms of a clinical term through Biomedical entity linking. With the introduction of transformer-based language models, there has been significant progress in Biomedical entity linking. In this work, we focus on learning through synonym pairs associated with the entities. As compared to the existing approaches, our approach significantly reduces the training data and resource consumption. Moreover, we propose a suite of context-based and context-less reranking techniques for performing the entity disambiguation. Overall, we achieve similar performance to the state-of-the-art zero-shot and distant supervised entity linking techniques on the Medmentions dataset, the largest annotated dataset on UMLS, without any domain-based training. Finally, we show that retrieval performance alone might not be sufficient as an evaluation metric and introduce an article level quantitative and qualitative analysis to reveal further insights on the performance of entity linking methods.
翻译:临床文本富含信息,包含治疗、药物和解剖学等多种临床术语的提及。多个术语可能指向相同的核心概念,这些概念可被称为临床实体。统一医学语言系统(UMLS)等本体被开发并维护,以存储数百万个临床实体,包括定义、关系及其他相关信息。这些本体通过生物医学实体链接将临床术语的不同表层形式归一化,从而实现临床文本的标准化。随着基于Transformer的语言模型的引入,生物医学实体链接领域取得了显著进展。本研究侧重于通过实体相关的同义词对进行学习。与现有方法相比,我们的方法显著减少了训练数据和资源消耗。此外,我们提出了一套基于上下文和无上下文的重新排序技术,用于执行实体消歧。总体而言,我们在Medmentions数据集(UMLS上最大的标注数据集)上实现了与最先进的零样本和远程监督实体链接技术相当的性能,且无需任何领域特定训练。最后,我们表明仅靠检索性能可能不足以作为评估指标,并引入了文章级别的定量和定性分析,以进一步揭示实体链接方法性能的深层洞察。