Learning vectors that capture the meaning of concepts remains a fundamental challenge. Somewhat surprisingly, perhaps, pre-trained language models have thus far only enabled modest improvements to the quality of such concept embeddings. Current strategies for using language models typically represent a concept by averaging the contextualised representations of its mentions in some corpus. This is potentially sub-optimal for at least two reasons. First, contextualised word vectors have an unusual geometry, which hampers downstream tasks. Second, concept embeddings should capture the semantic properties of concepts, whereas contextualised word vectors are also affected by other factors. To address these issues, we propose two contrastive learning strategies, based on the view that whenever two sentences reveal similar properties, the corresponding contextualised vectors should also be similar. One strategy is fully unsupervised, estimating the properties which are expressed in a sentence from the neighbourhood structure of the contextualised word embeddings. The second strategy instead relies on a distant supervision signal from ConceptNet. Our experimental results show that the resulting vectors substantially outperform existing concept embeddings in predicting the semantic properties of concepts, with the ConceptNet-based strategy achieving the best results. These findings are furthermore confirmed in a clustering task and in the downstream task of ontology completion.
翻译:学习能够捕捉概念含义的向量仍然是一个基本挑战。令人有些意外的是,预训练语言模型迄今为止对这类概念嵌入质量的改进仅限于微小的提升。当前使用语言模型的策略通常通过平均概念在某个语料库中出现的上下文表示来表示该概念。这种策略可能至少因两个原因而次优。首先,上下文词向量具有不寻常的几何结构,这会阻碍下游任务。其次,概念嵌入应捕捉概念的语义属性,而上下文词向量还受到其他因素的影响。为解决这些问题,我们提出了两种对比学习策略,基于以下观点:每当两个句子揭示相似的属性时,对应的上下文向量也应该是相似的。一种策略是完全无监督的,从上下文词嵌入的邻域结构中估计句子中所表达的属性。另一种策略则依赖于来自ConceptNet的远程监督信号。我们的实验结果表明,所得到的向量在预测概念语义属性方面显著优于现有的概念嵌入,其中基于ConceptNet的策略取得了最佳结果。这些发现还在聚类任务和本体补全下游任务中得到了进一步证实。