When large language models encounter conflicting information in context, which memories survive -- early or recent? We adapt classical interference paradigms from cognitive psychology to answer this question, testing 39 LLMs across diverse architectures and scales. Every model shows the same pattern: proactive interference (PI) dominates retroactive interference (RI) universally (Cohen's d = 1.73, p < 0.0001), meaning early encodings are protected at the cost of recent information -- the opposite of human memory, where RI typically dominates. Three findings indicate that RI and PI reflect separate memory mechanisms. RI and PI are uncorrelated (R^2 = 0.044), rejecting a unified "memory capacity." Model size predicts RI resistance (R^2 = 0.49) but not PI (R^2 = 0.06, n.s.) -- only RI is capacity-dependent. And error analysis reveals distinct failure modes: RI failures are passive retrieval failures (51%), while PI failures show active primacy intrusion (56%); both show <1% hallucination. These patterns parallel the consolidation-retrieval distinction in cognitive science, suggesting that transformer attention creates a primacy bias with direct implications for interference-heavy applications.
翻译:当大语言模型在上下文中遇到冲突信息时,哪些记忆得以保留——早期的还是近期的?我们借鉴认知心理学中的经典干扰范式来解答此问题,对涵盖不同架构和规模的39个大语言模型进行了测试。所有模型均呈现相同模式:前摄干扰普遍主导倒摄干扰(Cohen's d = 1.73,p < 0.0001),即早期编码受到保护而牺牲近期信息——这与人类记忆通常由倒摄干扰主导的现象相反。三项发现表明,前摄干扰与倒摄干扰反映不同的记忆机制:二者无相关性(R^2 = 0.044),否定了统一的“记忆容量”假设;模型规模能预测倒摄干扰抗性(R^2 = 0.49)却无法预测前摄干扰(R^2 = 0.06,不显著)——仅倒摄干扰具有容量依赖性;误差分析揭示了不同的失败模式:倒摄干扰失败表现为被动检索失败(51%),而前摄干扰失败表现为主动首因侵入(56%),两者幻觉率均低于1%。这些模式与认知科学中的巩固-检索区分相呼应,表明Transformer注意力机制产生的首因偏差对干扰密集型应用具有直接影响。