We investigate whether methods of human mathematics pedagogy can guide the training of language models toward arithmetic reasoning. Building on the GASING method -- an Indonesian pedagogy that solves basic arithmetic through a left-to-right procedure aligned with the causal order of token generation -- we operationalize each operation as a computational procedure whose execution trace is serialized into natural-language Chain-of-Thought (CoT) supervision. A small GPT-2 decoder (86M parameters) with a syllabic-agglutinative TOBA tokenizer for Indonesian is trained from scratch on this data using only a next-token prediction objective, without reinforcement learning or reward-based optimization. Monitoring training reveals three distinct learning phases, and mechanistic analyses -- attention-masking interventions on the CoT information graph, residual-stream probing, and logit-lens inspection -- show that the model first internalizes a procedural pathway and subsequently develops an associative, ``mental-arithmetic'' capacity that retrieves intermediate results without explicit step-by-step computation. The trained model reaches over 80% accuracy on held-out problems and attains competitive performance against substantially larger language models, indicating that targeted, pedagogically grounded training can yield strong and economical arithmetic capability at small scale.
翻译:我们探究人类数学教学法能否引导语言模型训练中的算术推理能力。基于GASING方法——一种通过从左到右过程(与因果顺序的token生成对齐)解决基础算术的印尼教学法——我们将每个运算操作具体化为计算程序,并将其执行轨迹序列化为自然语言的思维链(CoT)监督。使用仅含下一个token预测目标的训练目标,从零训练一个带有音节-黏着式TOBA分词器(用于印尼语)的小型GPT-2解码器(86M参数),无需强化学习或基于奖励的优化。监控训练过程揭示了三个不同的学习阶段,机制分析——包括对CoT信息图的注意力遮蔽干预、残差流探测以及logit透镜检查——表明模型首先内化程序化路径,随后发展出联想式“心算”能力,能够无需显式逐步计算即可检索中间结果。训练后的模型在保留问题上准确率超过80%,并达到与规模更大的语言模型竞争力相当的表现,表明基于教学法的定向训练能在小规模下产生强大且经济的算术能力。