We introduce The Benchmark of Linguistic Minimal Pairs (shortened to BLiMP), a challenge set for evaluating what language models (LMs) know about major grammatical phenomena in English. BLiMP consists of 67 sub-datasets, each containing 1000 minimal pairs isolating specific contrasts in syntax, morphology, or semantics. The data is automatically generated according to expert-crafted grammars, and aggregate human agreement with the labels is 96.4%. We use it to evaluate n-gram, LSTM, and Transformer (GPT-2 and Transformer-XL) LMs. We find that state-of-the-art models identify morphological contrasts reliably, but they struggle with semantic restrictions on the distribution of quantifiers and negative polarity items and subtle syntactic phenomena such as extraction islands.
翻译:我们提出了英语语言最小对比对基准测试(简称BLiMP),这是一个用于评估语言模型对英语主要语法现象掌握程度的挑战性数据集。BLiMP包含67个子数据集,每个子数据集包含1000个针对句法、形态或语义特定对比关系的的最小对比对。这些数据根据专家编写的语法规则自动生成,人工标注整体一致性达到96.4%。我们利用该基准测试评估了n-gram、LSTM以及Transformer(GPT-2和Transformer-XL)语言模型。研究发现,当前最优模型能够可靠识别形态对比,但在处理量词与负极词分布中的语义限制以及提取岛等微妙句法现象时仍存在困难。