Building natural language inference (NLI) benchmarks that are both challenging for modern techniques, and free from shortcut biases is difficult. Chief among these biases is "single sentence label leakage," where annotator-introduced spurious correlations yield datasets where the logical relation between (premise, hypothesis) pairs can be accurately predicted from only a single sentence, something that should in principle be impossible. We demonstrate that despite efforts to reduce this leakage, it persists in modern datasets that have been introduced since its 2018 discovery. To enable future amelioration efforts, introduce a novel model-driven technique, the progressive evaluation of cluster outliers (PECO) which enables both the objective measurement of leakage, and the automated detection of subpopulations in the data which maximally exhibit it.
翻译:构建既对现代技术具有挑战性、又免于捷径偏差的自然语言推理基准数据集是困难的。其中最主要的偏差是“单句标签泄露”,即标注者引入的伪相关导致数据集中(前提、假设)对之间的逻辑关系仅通过单一句子即可准确预测——而这在原则上本不可能。我们证明,尽管已有减少这种泄露的努力,但自2018年发现以来,在现代数据集中该问题依然存在。为了推动未来的改进工作,我们提出了一种新颖的模型驱动技术——簇异常值的渐进评估(PECO),该技术既能实现泄露的客观量化,又能自动检测数据中最大程度呈现泄露的子群体。