Mathematical reasoning in large language models (LLMs) has garnered attention in recent research, but there is limited understanding of how these models process and store information related to arithmetic tasks. In this paper, we present a mechanistic interpretation of LLMs for arithmetic-based questions using a causal mediation analysis framework. By intervening on the activations of specific model components and measuring the resulting changes in predicted probabilities, we identify the subset of parameters responsible for specific predictions. We analyze two pre-trained language models with different sizes (2.8B and 6B parameters). Experimental results reveal that a small set of mid-late layers significantly affect predictions for arithmetic-based questions, with distinct activation patterns for correct and wrong predictions. We also investigate the role of the attention mechanism and compare the model's activation patterns for arithmetic queries with the prediction of factual knowledge. Our findings provide insights into the mechanistic interpretation of LLMs for arithmetic tasks and highlight the specific components involved in arithmetic reasoning.
翻译:大型语言模型(LLMs)中的数学推理在近期研究中备受关注,但关于这些模型如何处理和存储与算术任务相关的信息,目前理解有限。本文基于因果中介分析框架,对面向算术问题的LLMs进行机制性解释。通过干预特定模型组件的激活状态并测量预测概率的变化,我们识别出负责特定预测的参数子集。我们分析了两个不同规模(28亿和60亿参数)的预训练语言模型。实验结果表明,少量中后层对算术问题的预测具有显著影响,且正确预测与错误预测呈现不同的激活模式。我们还探究了注意机制的作用,并将模型对算术查询的激活模式与事实知识预测进行对比。本研究为理解LLMs在算术任务中的机制性解释提供了新视角,并揭示了参与算术推理的特定组件。