In today's technologically driven world, the rapid spread of fake news, particularly during critical events like elections, poses a growing threat to the integrity of information. To tackle this challenge head-on, we introduce FakeWatch, a comprehensive framework carefully designed to detect fake news. Leveraging a newly curated dataset of North American election-related news articles, we construct robust classification models. Our framework integrates a model hub comprising of both traditional machine learning (ML) techniques and cutting-edge Language Models (LMs) to discern fake news effectively. Our overarching objective is to provide the research community with adaptable and precise classification models adept at identifying the ever-evolving landscape of misinformation. Quantitative evaluations of fake news classifiers on our dataset reveal that, while state-of-the-art LMs exhibit a slight edge over traditional ML models, classical models remain competitive due to their balance of accuracy and computational efficiency. Additionally, qualitative analyses shed light on patterns within fake news articles. This research lays the groundwork for future endeavors aimed at combating misinformation, particularly concerning electoral processes. We provide our labeled data and model publicly for use and reproducibility.
翻译:在当今技术驱动的世界中,假新闻的快速传播(尤其是在选举等关键事件期间)对信息完整性构成了日益严重的威胁。为直击这一挑战,我们提出了FakeWatch——一个精心设计的综合假新闻检测框架。通过利用新整理的北美选举相关新闻文章数据集,我们构建了稳健的分类模型。该框架集成了包含传统机器学习技术与前沿语言模型的模型中心,以有效识别假新闻。我们的总体目标是为研究社区提供适应性强且精确的分类模型,使其能够精准识别不断演变的虚假信息生态。基于该数据集对假新闻分类器进行的定量评估表明,尽管前沿语言模型略优于传统机器学习模型,但经典模型凭借其准确性与计算效率的平衡仍具竞争力。此外,定性分析揭示了假新闻文章中的内在规律。本研究为未来打击虚假信息(尤其是涉及选举过程)的举措奠定了基础。我们公开提供标注数据与模型,以供使用及复现验证。