Large language models (LLMs) are increasingly used to generate scientific reports, but they can produce references that appear plausible while containing corrupted metadata or pointing to papers that do not exist. We introduce CiteCheck, a hybrid framework for citation hallucination detection that verifies whether a citation corresponds to a real scholarly work and whether its metadata is faithful to that work. CiteCheck retrieves candidate publications from external scholarly sources, compares the citation against the retrieved candidate using a structured LLM verifier, and maps verifier scores into three labels: Exact, Minor, and Major. We also construct a 982-citation physics benchmark with controlled corruptions that capture both subtle metadata drift and fully fabricated references. On the held-out test set, CiteCheck achieves 88.7 macro-F1 and 88.9% accuracy, outperforming GPT, Claude, and Gemini baselines, including web-search and few-shot variants. These results show that reliable citation verification benefits from combining scholarly retrieval, structured LLM-based comparison, and calibrated decision rules.
翻译:摘要:大语言模型(LLMs)日益被用于生成科学报告,但可能生成看似合理却包含错误元数据或指向不存在论文的参考文献。我们提出CiteCheck这一混合框架,通过验证引用是否对应真实学术著作及其元数据是否忠实于该著作来检测引用幻觉。该框架从外部学术资源检索候选出版物,利用结构化LLM验证器将引用与候选结果进行比对,并将验证器评分映射为三类标签:精确、微小错误和重大错误。我们还构建了包含982条引用的物理学基准测试集,其中包含控制性错误,同时涵盖细微元数据偏差与完全捏造的参考文献。在保留测试集上,CiteCheck实现了88.7的宏F1值和88.9%的准确率,优于包括网络搜索和少样本变体在内的GPT、Claude和Gemini基线模型。这些结果表明,结合学术检索、结构化LLM比对与校准决策规则能够有效提升引用验证的可靠性。