This paper presents a partial reproduction of Generating Fact Checking Explanations by Anatanasova et al (2020) as part of the ReproHum element of the ReproNLP shared task to reproduce the findings of NLP research regarding human evaluation. This shared task aims to investigate the extent to which NLP as a field is becoming more or less reproducible over time. Following the instructions provided by the task organisers and the original authors, we collect relative rankings of 3 fact-checking explanations (comprising a gold standard and the outputs of 2 models) for 40 inputs on the criteria of Coverage. The results of our reproduction and reanalysis of the original work's raw results lend support to the original findings, with similar patterns seen between the original work and our reproduction. Whilst we observe slight variation from the original results, our findings support the main conclusions drawn by the original authors pertaining to the efficacy of their proposed models.
翻译:本文作为ReproNLP共享任务中ReproHum环节的一部分,对Anatanasova等人(2020)提出的《生成事实验证解释》一文进行了部分复现,旨在再现NLP研究中关于人类评估的发现。该共享任务致力于探究NLP领域随时间推移在可复现性方面的变化趋势。遵循任务组织者及原始作者提供的指导,我们针对40个输入样本,基于覆盖度标准收集了3种事实验证解释(包括一个金标准和两个模型输出)的相对排序。我们的复现结果及对原始工作原始数据的再分析支持了原始发现,并在原始研究与我们复现的结果之间观察到相似模式。尽管我们注意到与原始结果存在细微差异,但我们的发现支持原始作者就其提出模型有效性得出的主要结论。