Measuring semantic change has thus far remained a task where methods using contextual embeddings have struggled to improve upon simpler techniques relying only on static word vectors. Moreover, many of the previously proposed approaches suffer from downsides related to scalability and ease of interpretation. We present a simplified approach to measuring semantic change using contextual embeddings, relying only on the most probable substitutes for masked terms. Not only is this approach directly interpretable, it is also far more efficient in terms of storage, achieves superior average performance across the most frequently cited datasets for this task, and allows for more nuanced investigation of change than is possible with static word vectors.
翻译:迄今为止,语义变化测量领域仍存在一个现象:依赖上下文嵌入的方法始终难以超越仅使用静态词向量的简单技术。此外,先前提出的许多方法都受限于可扩展性和解释便捷性等问题。本文提出一种简化的语义变化测量方法,该方法利用上下文嵌入,仅依赖掩码词项的最高概率替代词。这不仅使该方法具有直接可解释性,而且在存储效率上显著提升,在语义变化检测领域最常引用的数据集中实现了更优的平均性能表现,同时能够比静态词向量更深入地探究语义变化细节。