Retrieval-Augmented Generation (RAG) models are critically undermined by citation hallucinations, a deceptive failure where a model cites a source that fails to support its claim. While existing work attributes hallucination to a simple over-reliance on parametric knowledge, we reframe this failure as an evolving, scale-dependent coordination failure between the Attention (reading) and Feed-Forward Network (recalling) pathways. We introduce FACTUM (Framework for Attesting Citation Trustworthiness via Underlying Mechanisms), a framework of four mechanistic scores: Contextual Alignment (CAS), Attention Sink Usage (BAS), Parametric Force (PFS), and Pathway Alignment (PAS). Our analysis reveals that correct citations are consistently marked by higher parametric force (PFS) and greater use of the attention sink (BAS) for information synthesis. Crucially, we find that "one-size-fits-all" theories are insufficient as the signature of correctness evolves with scale: while the 3B model relies on high pathway alignment (PAS), our best-performing 8B detector identifies a shift toward a specialized strategy where pathways provide distinct, orthogonal information. By capturing this complex interplay, FACTUM outperforms state-of-the-art baselines by up to 37.5% in AUC. Our results demonstrate that high parametric force is constructive when successfully coordinated with the Attention pathway, paving the way for more nuanced and reliable RAG systems.
翻译:检索增强生成(RAG)模型因引用幻觉而受到严重削弱——这是一种欺骗性故障,即模型引用了一个无法支撑其主张的源文献。现有研究将幻觉归因于对参数化知识的简单过度依赖,但我们重新将此故障定义为注意力(阅读)通路与前馈网络(回忆)通路之间一种演化性的、随规模尺度而变化的协调失效。我们提出FACTUM(基于潜在机制的引用可信度验证框架),该框架包含四项机制评分:上下文对齐度(CAS)、注意力汇聚使用度(BAS)、参数化驱动力(PFS)和通路对齐度(PAS)。分析表明,正确引用始终伴随着更高的参数化驱动力(PFS)和更充分的注意力汇聚(BAS)用于信息整合。关键的是,我们发现了"一刀切"理论的不足——正确引用的特征随规模尺度而演变:3B参数模型依赖高通路对齐度(PAS),而我们表现最佳的8B参数检测器识别出一种向专门化策略的转变,即通路提供不同的正交信息。通过捕捉这种复杂交互作用,FACTUM在AUC指标上较现有最优基线模型提升高达37.5%。我们的结果表明:当高参数化驱动力成功与注意力通路协调时具有建设性,这为构建更精细、更可靠的RAG系统铺平了道路。