We present PaRTE, a collection of 1,126 pairs of Recognizing Textual Entailment (RTE) examples to evaluate whether models are robust to paraphrasing. We posit that if RTE models understand language, their predictions should be consistent across inputs that share the same meaning. We use the evaluation set to determine if RTE models' predictions change when examples are paraphrased. In our experiments, contemporary models change their predictions on 8-16\% of paraphrased examples, indicating that there is still room for improvement.
翻译:我们提出了PaRTE,一个包含1126对文本蕴含识别(RTE)示例的数据集,用于评估模型对释义的鲁棒性。我们认为,如果RTE模型理解语言,其预测应在语义相同的输入间保持一致性。通过该评估集,我们检测RTE模型在示例被释义后预测是否发生变化。实验表明,当代模型在8-16%的释义示例上改变了预测结果,这表明仍存在改进空间。