Automated math word problem solvers based on neural networks have successfully managed to obtain 70-80\% accuracy in solving arithmetic word problems. However, it has been shown that these solvers may rely on superficial patterns to obtain their equations. In order to determine what information math word problem solvers use to generate solutions, we remove parts of the input and measure the model's performance on the perturbed dataset. Our results show that the model is not sensitive to the removal of many words from the input and can still manage to find a correct answer when given a nonsense question. This indicates that automatic solvers do not follow the semantic logic of math word problems, and may be overfitting to the presence of specific words.
翻译:基于神经网络的自动数学应用题求解器已成功在算术应用题中实现了70-80%的准确率。然而,研究表明这些求解器可能依赖表面模式来获取其方程式。为了确定数学应用题求解器在生成解答时使用了哪些信息,我们移除了输入的部分内容,并在扰动后的数据集上测量模型性能。我们的结果表明,模型对输入中许多词语的移除不敏感,并且在面对无意义问题时仍能设法找到正确答案。这表明自动求解器并未遵循数学应用题的语义逻辑,可能存在对特定词语出现的过拟合现象。