Irrelevant numbers in a prompt can shift language model judgments, producing anchoring effects in numerical reasoning. We study where this anchor-sensitive signal is carried inside language models using a controlled multiple-choice setup with shared answer options. We define a logit-difference metric comparing the correct answer option with the answer option corresponding to the anchor, and validate that it tracks behavioral anchoring. Using attribution-based circuit localization on 7B--8B Qwen and Llama base and instruction-tuned models, we find that edge-level methods recover this signal more faithfully than node-level methods. Low- and high-anchor circuits transfer strongly within a model, suggesting shared pathway structure across anchor direction. However, sparse transfer across base and instruction-tuned variants is less reliable, indicating that post-training changes which pathways matter most. Overall, our results provide a mechanistic account of how anchoring-related decision signals are carried inside language models.
翻译:提示词中的无关数字会改变语言模型的判断,从而在数值推理中产生锚定效应。我们通过一种具有共享答案选项的受控多选题设置,研究语言模型内部承载这种锚定敏感信号的路径。我们定义了一个对数几率差指标,用于比较正确答案选项与锚定对应的答案选项,并验证其能追踪行为锚定。利用基于归因的电路定位方法,在7B-8B的Qwen和Llama基础模型及指令微调模型上,我们发现边级方法比节点级方法更能忠实恢复该信号。低锚定与高锚定电路在模型内具有很强的迁移性,表明锚定方向存在共享路径结构。然而,在基础模型与指令微调变体之间的稀疏迁移可靠性较低,这表明后训练过程改变了哪些路径最为关键。总体而言,我们的结果为锚定相关决策信号在语言模型内部的承载机制提供了机理层面的解释。