Previous studies have typically assumed that large language models are unable to accurately perform arithmetic operations, particularly multiplication of >8 digits, and operations involving decimals and fractions, without the use of calculator tools. This paper aims to challenge this misconception. With sufficient training data, a 2 billion-parameter language model can accurately perform multi-digit arithmetic operations with almost 100% accuracy without data leakage, significantly surpassing GPT-4 (whose multi-digit multiplication accuracy is only 4.3%). We also demonstrate that our MathGLM, fine-tuned from GLM-10B on a dataset with additional multi-step arithmetic operations and math problems described in text, achieves similar performance to GPT-4 on a 5,000-samples Chinese math problem test set.
翻译:以往研究通常假设大语言模型在没有计算器工具的情况下,无法准确执行算术运算,尤其是超过8位数的乘法以及涉及小数和分数的运算。本文旨在挑战这一误解。通过充足的训练数据,一个20亿参数的语言模型可以在几乎100%准确率的情况下执行多位算术运算,且不存在数据泄露问题,显著超越GPT-4(其多位乘法准确率仅为4.3%)。我们还证明了,我们基于包含额外多步算术运算和文本描述的数学问题的数据集,从GLM-10B微调得到的MathGLM,在一个包含5,000个样本的中文数学问题测试集上,取得了与GPT-4相当的性能。