Automatic fake news detection with machine learning can prevent the dissemination of false statements before they gain many views. Several datasets labeling statements as legitimate or false have been created since the 2016 United States presidential election for the prospect of training machine learning models. We evaluate the robustness of both traditional and deep state-of-the-art models to gauge how well they may perform in the real world. We find that traditional models tend to generalize better to data outside the distribution it was trained on compared to more recently-developed large language models, though the best model to use may depend on the specific task at hand.
翻译:基于机器学习的自动假新闻检测能够在虚假言论获得大量关注之前阻止其传播。自2016年美国总统大选以来,为训练机器学习模型,已有多个标注语句真实性(真实或虚假)的数据集被创建。我们评估了传统模型与最先进的深度学习模型的鲁棒性,以衡量它们在现实世界中的实际表现。研究发现,与近期开发的大型语言模型相比,传统模型通常能更好地泛化到训练数据分布之外的样本中;然而,最佳模型的选择仍可能取决于具体任务的需求。