Math word problems are critical K-8 educational tools, but writing them is time-consuming and requires domain expertise. We suggest that language models can support K-8 math education by automatically generating problems at scale. To be educational, generated problems must be 1) solvable, 2) accurate, and 3) appropriate. Existing datasets are unlabeled for these criteria, making them ill-suited for training problem generators. We introduce MATHWELL, a Llama-2 (70B) model iteratively finetuned to generate K-8 math word problems using data from expert annotation. Using MATHWELL, we generate the largest English word problem dataset with Program of Thought (PoT) rationales to date, containing 20,490 problems. 3,484 are scored by domain experts who find MATHWELL has a 40% higher share of problems that have executable solutions and meet all criteria than alternatives, with 74% of its problems with executable solutions being solvable, accurate, and appropriate. We release our model, data, and annotations.
翻译:数学文字题是K-8阶段的关键教育工具,但编写这类题目耗时且需要专业知识。我们提出,语言模型可通过自动生成大规模题目来支持K-8数学教育。为确保教育价值,生成的题目必须满足:1)可求解性、2)准确性、3)适切性。现有数据集缺乏对这些标准的标注,因此不适合训练题目生成模型。我们提出了MATHWELL——基于Llama-2 (70B)模型、通过专家标注数据迭代微调的K-8数学文字题生成模型。利用MATHWELL,我们生成了迄今为止规模最大、包含思维程序(Program of Thought, PoT)推导过程的英语文字题数据集,共计20,490道题目。其中3,484道题目经领域专家评分,结果显示:MATHWELL生成的可执行方案题目中,满足所有标准的比例比替代方案高出40%,且74%具有可执行方案的题目兼具可求解性、准确性和适切性。我们将公开模型、数据集及标注结果。