Seeking health-related advice on the internet has become a common practice in the digital era. Determining the trustworthiness of medical claims found online and finding appropriate evidence for this information is increasingly challenging. Fact-checking has emerged as an approach to assess the veracity of factual claims using evidence from credible knowledge sources. To help advance the automation of this task, in this paper, we introduce a novel dataset of 750 health-related claims, labeled for veracity by medical experts and backed with evidence from appropriate clinical studies. We provide an analysis of the dataset, highlighting its characteristics and challenges. The dataset can be used for Machine Learning tasks related to automated fact-checking such as evidence retrieval, veracity prediction, and explanation generation. For this purpose, we provide baseline models based on different approaches, examine their performance, and discuss the findings.
翻译:在数字时代,通过网络寻求健康建议已成为普遍做法。判断网上医学声明的可信度并为其寻找适当证据日益具有挑战性。事实核查作为一种方法应运而生,它通过可信知识来源中的证据来评估事实性声明的真实性。为促进此类任务的自动化进程,本文提出了一个包含750条健康相关声明的新数据集,这些声明由医学专家标注真实性,并附有来自适当临床研究的证据支持。我们对数据集进行了分析,强调了其特征与挑战。该数据集可用于与自动事实核查相关的机器学习任务,如证据检索、真实性预测和解释生成。为此,我们提供了基于不同方法的基线模型,检验了它们的性能,并讨论了相关发现。