The increasing rate at which scientific knowledge is discovered and health claims shared online has highlighted the importance of developing efficient fact-checking systems for scientific claims. The usual setting for this task in the literature assumes that the documents containing the evidence for claims are already provided and annotated or contained in a limited corpus. This renders the systems unrealistic for real-world settings where knowledge sources with potentially millions of documents need to be queried to find relevant evidence. In this paper, we perform an array of experiments to test the performance of open-domain claim verification systems. We test the final verdict prediction of systems on four datasets of biomedical and health claims in different settings. While keeping the pipeline's evidence selection and verdict prediction parts constant, document retrieval is performed over three common knowledge sources (PubMed, Wikipedia, Google) and using two different information retrieval techniques. We show that PubMed works better with specialized biomedical claims, while Wikipedia is more suited for everyday health concerns. Likewise, BM25 excels in retrieval precision, while semantic search in recall of relevant evidence. We discuss the results, outline frequent retrieval patterns and challenges, and provide promising future directions.
翻译:科学知识的发现速度不断加快,健康声明在网络上的传播日益频繁,这凸显了开发高效科学声明事实核查系统的必要性。文献中对此任务的通常设定是,已提供或标注了包含证据的文档,或将其限定在小型语料库中。这使得系统在现实场景中不切实际——需要通过查询包含数百万文档的知识源来寻找相关证据。本文通过一系列实验,测试了开放域声明验证系统的性能。我们在四个生物医学与健康声明的数据集上,在不同设定下对系统的最终裁决预测进行了测试。在保持流水线的证据选择与裁决预测部分不变的情况下,我们使用两种不同的信息检索技术,从三种常见知识源(PubMed、Wikipedia、Google)中进行文档检索。结果表明:PubMed更适合处理专业的生物医学声明,而Wikipedia则更适用于日常健康问题。同样,BM25在检索精确度上表现优异,而语义搜索在相关证据的召回率上更胜一筹。我们讨论了实验结果,总结了常见的检索模式与挑战,并提出了有前景的未来研究方向。