Fact checking can be an effective strategy against misinformation, but its implementation at scale is impeded by the overwhelming volume of information online. Recent artificial intelligence (AI) language models have shown impressive ability in fact-checking tasks, but how humans interact with fact-checking information provided by these models is unclear. Here, we investigate the impact of fact-checking information generated by a popular large language model (LLM) on belief in, and sharing intent of, political news headlines in a preregistered randomized control experiment. Although the LLM accurately identifies most false headlines (90%), we find that this information does not significantly improve participants' ability to discern headline accuracy or share accurate news. In contrast, viewing human-generated fact checks enhances discernment in both cases. Subsequent analysis reveals that the AI fact-checker is harmful in specific cases: it decreases beliefs in true headlines that it mislabels as false and increases beliefs in false headlines that it is unsure about. On the positive side, AI fact-checking information increases the sharing intent for correctly labeled true headlines. When participants are given the option to view LLM fact checks and choose to do so, they are significantly more likely to share both true and false news but only more likely to believe false headlines. Our findings highlight an important source of potential harm stemming from AI applications and underscore the critical need for policies to prevent or mitigate such unintended consequences.
翻译:事实核查是对抗错误信息的有效策略,但其大规模实施受限于线上信息的海量规模。近期人工智能语言模型在事实核查任务中展现出卓越能力,但人类如何与这些模型提供的事实核查信息互动尚不明确。本研究通过预注册随机对照实验,探究了流行大型语言模型生成的事实核查信息对政治新闻标题可信度判断与分享意愿的影响。尽管该模型能准确识别大部分虚假标题(90%),但我们发现这些信息并未显著提升参与者辨别标题准确性或分享真实新闻的能力。相比之下,查看人工生成的事实核查则在两方面均能提升辨别力。后续分析表明,AI事实核查器在特定情况下存在危害:它会降低被误标为虚假的真实标题可信度,并提升其无法确定的虚假标题可信度。从积极方面看,AI事实核查信息能提升被正确标注的真实标题的分享意愿。当参与者获得查看大型语言模型事实核查的选项并选择使用时,他们分享真实与虚假新闻的可能性均显著增加,但仅对虚假标题的可信度判断出现提升。我们的研究结果揭示了人工智能应用潜在危害的重要来源,并强调亟需制定政策以预防或减轻此类非预期后果。