This paper shines a light on the potential of definition-based semantic models for detecting idiomatic and semi-idiomatic multiword expressions (MWEs) in clinical terminology. Our study focuses on biomedical entities defined in the UMLS ontology and aims to help prioritize the translation efforts of these entities. In particular, we develop an effective tool for scoring the idiomaticity of biomedical MWEs based on the degree of similarity between the semantic representations of those MWEs and a weighted average of the representation of their constituents. We achieve this using a biomedical language model trained to produce similar representations for entity names and their definitions, called BioLORD. The importance of this definition-based approach is highlighted by comparing the BioLORD model to two other state-of-the-art biomedical language models based on Transformer: SapBERT and CODER. Our results show that the BioLORD model has a strong ability to identify idiomatic MWEs, not replicated in other models. Our corpus-free idiomaticity estimation helps ontology translators to focus on more challenging MWEs.
翻译:本文揭示了基于定义的语义模型在检测临床术语中习语性和半习语性多词表达(MWEs)方面的潜力。我们的研究聚焦于UMLS本体中定义的生物医学实体,旨在帮助优先处理这些实体的翻译工作。具体而言,我们开发了一个有效工具,根据这些多词表达的语义表示与其成分表示加权平均值之间的相似度,对生物医学多词表达的习语性进行评分。我们通过使用一个名为BioLORD的生物医学语言模型来实现这一目标,该模型经过训练,使得实体名称及其定义能产生相似的表示。通过将BioLORD模型与另外两种基于Transformer的最新生物医学语言模型SapBERT和CODER进行比较,突显了这种基于定义方法的重要性。我们的结果表明,BioLORD模型具有很强的识别习语性多词表达的能力,这是其他模型所不具备的。我们这种无需语料库的习语性评估方法有助于本体翻译人员将精力集中在更具挑战性的多词表达上。