We present a human-in-the-loop evaluation framework for fact-checking novel misinformation claims and identifying social media messages that support them. Our approach extracts check-worthy claims, which are aggregated and ranked for review. Stance classifiers are then used to identify tweets supporting novel misinformation claims, which are further reviewed to determine whether they violate relevant policies. To demonstrate the feasibility of our approach, we develop a baseline system based on modern NLP methods for human-in-the-loop fact-checking in the domain of COVID-19 treatments. Using our baseline system, we show that human fact-checkers can identify 124 tweets per hour that violate Twitter's policies on COVID-19 misinformation. We will make our code, data, baseline models, and detailed annotation guidelines available to support the evaluation of human-in-the-loop systems that identify novel misinformation directly from raw user-generated content.
翻译:我们提出一种人机协同的评估框架,用于对新型虚假信息主张进行事实核查,并识别支持该主张的社交媒体信息。该方法首先提取值得核查的主张,经聚合排序后交由人工审核;随后利用立场分类器识别支持新型虚假信息主张的推文,并进一步评估其是否违反相关政策。为验证该方法的可行性,我们基于现代自然语言处理技术构建了针对COVID-19治疗领域的人机协同事实核查基线系统。实验表明,借助该基线系统,人工核查员每小时可识别124条违反Twitter关于COVID-19虚假信息政策的推文。我们将公开代码、数据、基线模型及详细标注指南,为直接对原始用户生成内容中的新型虚假信息进行鉴别人机协同评估提供支持。