In recent years, progress in NLU has been driven by benchmarks. These benchmarks are typically collected by crowdsourcing, where annotators write examples based on annotation instructions crafted by dataset creators. In this work, we hypothesize that annotators pick up on patterns in the crowdsourcing instructions, which bias them to write many similar examples that are then over-represented in the collected data. We study this form of bias, termed instruction bias, in 14 recent NLU benchmarks, showing that instruction examples often exhibit concrete patterns, which are propagated by crowdworkers to the collected data. This extends previous work (Geva et al., 2019) and raises a new concern of whether we are modeling the dataset creator's instructions, rather than the task. Through a series of experiments, we show that, indeed, instruction bias can lead to overestimation of model performance, and that models struggle to generalize beyond biases originating in the crowdsourcing instructions. We further analyze the influence of instruction bias in terms of pattern frequency and model size, and derive concrete recommendations for creating future NLU benchmarks.
翻译:近年来,自然语言理解领域的进展主要由各类基准测试驱动。这些基准测试通常通过众包方式收集数据,标注者根据数据集创建者设计的标注说明撰写示例。在本研究中,我们假设标注者会从众包说明中捕捉到某些模式,这些模式促使他们编写大量相似的示例,从而导致这些示例在收集的数据中被过度呈现。我们将这种形式的偏差称为“指令偏差”,并在14个近期的自然语言理解基准测试中对其进行了研究,结果表明指令示例通常表现出具体模式,而这些模式会被众包工作者传播到收集的数据中。这一发现拓展了先前的研究(Geva等人,2019),并提出了一个新的担忧:我们是否实际上是在对数据集创建者的指令进行建模,而非任务本身?通过一系列实验,我们证明指令偏差确实会导致模型性能被高估,并且模型难以泛化到超出众包说明中偏差来源的范围。我们进一步从模式频率和模型规模的角度分析了指令偏差的影响,并提出了创建未来自然语言理解基准测试的具体建议。