Large language models show promising capabilities for contextual fact-checking on social media: they can verify contested claims through deep research, synthesize evidence from multiple sources, and draft explanations at scale. However, prior work evaluates LLM fact-checking only in controlled settings using benchmarks or crowdworker judgments, leaving open how these systems perform in authentic platform environments. We present the first field evaluation of LLM-based fact-checking deployed on a live social media platform, testing performance directly through X Community Notes' AI writer feature over a three-month period. Our LLM writer, a multi-step pipeline that handles multimodal content (text, images, and videos), conducts web and platform-native search, and writes contextual notes, was deployed to write 1,614 notes on 1,597 tweets and compared against 1,332 human-written notes on the same tweets using 108,169 ratings from 42,521 raters. Direct comparison of note-level platform outcomes is complicated by differences in submission timing and rating exposure between LLM and human notes; we therefore pursue two complementary strategies: a rating-level analysis modeling individual rater evaluations, and a note-level analysis that equalizes rater exposure across note types. Rating-level analysis shows that LLM notes receive more positive ratings than human notes across raters with different political viewpoints, suggesting the potential for LLM-written notes to achieve the cross-partisan consensus. Note-level analysis confirms this advantage: among raters who evaluated all notes on the same post, LLM notes achieve significantly higher helpfulness scores. Our findings demonstrate that LLMs can contribute high-quality, broadly helpful fact-checking at scale, while highlighting that real-world evaluation requires careful attention to platform dynamics absent from controlled settings.


翻译:大型语言模型在社交媒体语境事实核查方面展现出可喜能力:它们可通过深度研究验证争议性言论,综合多源证据并规模化生成解释文本。然而,既有研究仅在受控环境中利用基准测试或众包评估考察LLM事实核查能力,尚未探明这些系统在真实平台环境中的表现。我们首次开展基于LLM的事实核查系统在实时社交媒体平台的现场评估——通过X平台社区笔记的AI撰写功能,在三个月周期内直接测试其性能。我们开发的LLM撰写器采用多阶段流水线架构,可处理多模态内容(文本、图像与视频)、执行网络及平台原生搜索、并撰写情境化注释。该工具共生成1,614条覆盖1,597条推文的注释,与人类撰写的1,332条同推文注释形成对照,共采用来自42,521名评价者的108,169份评分数据。由于LLM与人类注释在提交时序和评分曝光度上存在差异,直接进行注释级平台结果对比面临复杂性;我们因此采取两种互补策略:构建评估个体评价者评分的评分级分析模型,以及通过均衡注释类型间评价者曝光度的注释级分析。评分级分析显示,不同政治立场的评价者对LLM注释的正面评分均高于人类注释,表明LLM撰写的注释具备实现跨党派共识的潜力。注释级分析进一步验证此优势:在评估过同一帖子所有注释的评价者中,LLM注释获得显著更高的有用性评分。我们的研究证明,LLM能规模化生成高质量且具广泛实用性的事实核查建议,同时强调真实环境评估需特别关注受控环境中不存在的平台动态效应。

0
下载
关闭预览

相关内容

评估大语言模型在科学发现中的作用
专知会员服务
19+阅读 · 2025年12月19日
【斯坦福博士论文】大语言模型的AI辅助评估
专知会员服务
31+阅读 · 2025年3月30日
《使用生成式大语言模型进行多语言事件提取》最新85页
【CIKM2024】使用大型视觉语言模型的多模态虚假信息检测
生成型大型语言模型的自动事实核查:一项综述
专知会员服务
37+阅读 · 2024年7月6日
天大最新《大型语言模型评估》全面综述,111页pdf
专知会员服务
89+阅读 · 2023年10月31日
《利用 ChatGPT 实现高效事实核查》
专知会员服务
48+阅读 · 2023年10月25日
自然语言处理中的语言模型预训练方法
PaperWeekly
14+阅读 · 2018年10月21日
NLG ≠ 机器写作 | 专家专栏
量子位
13+阅读 · 2018年9月10日
TextInfoExp:自然语言处理相关实验(基于sougou数据集)
全球人工智能
12+阅读 · 2017年11月12日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
11+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
VIP会员
最新内容
边缘计算的军事应用
专知会员服务
2+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
4+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
10+阅读 · 8月7日
相关基金
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
11+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员