Automated fact-checking has drawn considerable attention over the past few decades due to the increase in the diffusion of misinformation on online platforms. This is often carried out as a sequence of tasks comprising (i) the detection of sentences circulating in online platforms which constitute claims needing verification, followed by (ii) the verification process of those claims. This survey focuses on the former, by discussing existing efforts towards detecting claims needing fact-checking, with a particular focus on multilingual data and methods. This is a challenging and fertile direction where existing methods are yet far from matching human performance due to the profoundly challenging nature of the issue. Especially, the dissemination of information across multiple social platforms, articulated in multiple languages and modalities demands more generalized solutions for combating misinformation. Focusing on multilingual misinformation, we present a comprehensive survey of existing multilingual claim detection research. We present state-of-the-art multilingual claim detection research categorized into three key factors of the problem, verifiability, priority, and similarity. Further, we present a detailed overview of the existing multilingual datasets along with the challenges and suggest possible future advancements.
翻译:自动化事实核查因在线平台虚假信息传播加剧,在过去数十年间受到广泛关注。该过程通常由两项任务构成:(i) 检测在线平台中需验证的声明性语句;(ii) 对这些声明进行验证。本文聚焦于前者,系统梳理了面向需事实核查声明的检测研究现状,特别关注多语言数据与方法。由于该问题本质极具挑战性,现有方法仍远未达到人类表现水平,这是充满难度的前沿方向。尤其当信息通过多语言及多模态形式在多个社交平台传播时,更需要通用化解决方案以应对虚假信息。本文基于多语言虚假信息视角,对现有声明检测研究进行综述,将当前最先进的多语言声明检测研究按可验证性、优先级与相似性三个关键维度进行归类。此外,本文详细梳理了现有多语言数据集及其面临的挑战,并对未来潜在发展方向提出建议。