Automated fact-checking systems verify claims against evidence to predict their veracity. In real-world scenarios, the retrieved evidence may not unambiguously support or refute the claim and yield conflicting but valid interpretations. Existing fact-checking datasets assume that the models developed with them predict a single veracity label for each claim, thus discouraging the handling of such ambiguity. To address this issue we present AmbiFC, a fact-checking dataset with 10k claims derived from real-world information needs. It contains fine-grained evidence annotations of 50k passages from 5k Wikipedia pages. We analyze the disagreements arising from ambiguity when comparing claims against evidence in AmbiFC, observing a strong correlation of annotator disagreement with linguistic phenomena such as underspecification and probabilistic reasoning. We develop models for predicting veracity handling this ambiguity via soft labels and find that a pipeline that learns the label distribution for sentence-level evidence selection and veracity prediction yields the best performance. We compare models trained on different subsets of AmbiFC and show that models trained on the ambiguous instances perform better when faced with the identified linguistic phenomena.
翻译:自动化事实验证系统通过对比证据来判定主张的真实性。但在现实场景中,检索到的证据可能无法明确支持或反驳主张,反而产生相互矛盾却合理解读的情况。现有事实验证数据集假设基于它们开发的模型会为每项主张预测单一真实性标签,从而阻碍了对这种含混性的处理。为解决该问题,我们提出AmbiFC——包含10,000条源自真实信息需求主张的事实验证数据集,该数据集对来自5,000个维基百科页面的50,000段文本进行了细粒度证据标注。我们分析了AmbiFC中主张与证据对比时因含混性产生的标注分歧,发现标注者分歧与语言现象(如欠指定性和概率推理)存在强相关性。我们开发了通过软标签处理含混性来预测真实性的模型,并发现采用句子级证据选择与真实性预测的标签分布学习流水线可获得最优性能。通过对比基于AmbiFC不同子集训练的模型,我们证明在面临上述语言现象时,基于含混实例训练的模型表现更优。