Metaphor identification aims at understanding whether a given expression is used figuratively in context. However, in this paper we show how existing metaphor identification datasets can be gamed by fully ignoring the potential metaphorical expression or the context in which it occurs. We test this hypothesis in a variety of datasets and settings, and show that metaphor identification systems based on language models without complete information can be competitive with those using the full context. This is due to the construction procedures to build such datasets, which introduce unwanted biases for positive and negative classes. Finally, we test the same hypothesis on datasets that are carefully sampled from natural corpora and where this bias is not present, making these datasets more challenging and reliable.
翻译:隐喻识别旨在判断给定表达在上下文中是否具有比喻性用法。然而,本文揭示了现有隐喻识别数据集可通过完全忽略潜在隐喻表达及其上下文而被轻易攻破。我们在多种数据集与实验设置中验证了这一假设,并表明基于语言模型的隐喻识别系统即使缺失完整信息,其表现仍可与利用完整上下文的系统相媲美。这一现象源于数据集构建流程中引入的针对正负类别的非预期偏差。最终,我们在从自然语料库中精心采样且不存在此类偏差的数据集上检验同一假设,证实此类数据集更具挑战性与可靠性。