Compositional generalization benchmarks seek to assess whether models can accurately compute meanings for novel sentences, but operationalize this in terms of logical form (LF) prediction. This raises the concern that semantically irrelevant details of the chosen LFs could shape model performance. We argue that this concern is realized for the COGS benchmark (Kim and Linzen, 2020). COGS poses generalization splits that appear impossible for present-day models, which could be taken as an indictment of those models. However, we show that the negative results trace to incidental features of COGS LFs. Converting these LFs to semantically equivalent ones and factoring out capabilities unrelated to semantic interpretation, we find that even baseline models get traction. A recent variable-free translation of COGS LFs suggests similar conclusions, but we observe this format is not semantically equivalent; it is incapable of accurately representing some COGS meanings. These findings inform our proposal for ReCOGS, a modified version of COGS that comes closer to assessing the target semantic capabilities while remaining very challenging. Overall, our results reaffirm the importance of compositional generalization and careful benchmark task design.
翻译:组合泛化基准旨在评估模型能否准确计算新句子的语义,但此类评估通常基于逻辑形式预测来实现。这引发了一个担忧:所选逻辑形式中与语义无关的细节可能影响模型性能。我们认为这一担忧在COGS基准(Kim和Linzen, 2020)中得到了证实。COGS提出了当前模型似乎无法解决的泛化任务划分,这可能导致对这些模型的负面评价。然而,我们证明了这些负面结果源于COGS逻辑形式的偶然特征。通过将这些逻辑形式转换为语义等价形式,并剥离与语义解释无关的能力,我们发现即使是基线模型也能取得进展。近期对COGS逻辑形式进行无变量转换的研究得出了类似结论,但我们观察到这种格式在语义上并不等价——它无法准确表示COGS的某些语义。这些发现推动了我们对ReCOGS的提出,这是COGS的改良版本,在保持高度挑战性的同时更贴近目标语义能力的评估。总体而言,我们的结果再次印证了组合泛化的重要性以及基准任务设计的严谨性。