In recent years, progress in NLU has been driven by benchmarks. These benchmarks are typically collected by crowdsourcing, where annotators write examples based on annotation instructions crafted by dataset creators. In this work, we hypothesize that annotators pick up on patterns in the crowdsourcing instructions, which bias them to write many similar examples that are then over-represented in the collected data. We study this form of bias, termed instruction bias, in 14 recent NLU benchmarks, showing that instruction examples often exhibit concrete patterns, which are propagated by crowdworkers to the collected data. This extends previous work (Geva et al., 2019) and raises a new concern of whether we are modeling the dataset creator's instructions, rather than the task. Through a series of experiments, we show that, indeed, instruction bias can lead to overestimation of model performance, and that models struggle to generalize beyond biases originating in the crowdsourcing instructions. We further analyze the influence of instruction bias in terms of pattern frequency and model size, and derive concrete recommendations for creating future NLU benchmarks.
翻译:近年来,自然语言理解(NLU)领域的进展主要由基准测试推动。这些基准测试通常通过众包方式收集,标注者根据数据集创建者精心设计的标注说明来撰写示例。在本研究中,我们假设标注者会捕捉众包说明中的模式,这些模式导致他们撰写大量相似的示例,从而在收集的数据中过度呈现。我们将这种形式的偏见称为“说明偏见”,并在14个近期NLU基准测试中进行了研究,结果表明,说明示例通常表现出具体的模式,这些模式被众包工作者传播到收集的数据中。这扩展了先前的研究(Geva等人,2019),并引发了一个新问题:我们是否在模拟数据集创建者的说明,而非任务本身。通过一系列实验,我们证明,说明偏见确实可能导致对模型性能的高估,且模型难以泛化到众包说明中起源的偏见之外。我们进一步从模式频率和模型规模的角度分析了说明偏见的影响,并为未来NLU基准测试的创建提出了具体建议。