Reasoning in mathematical domains remains a significant challenge for relatively small language models (LMs). Many current methods focus on specializing LMs in mathematical reasoning and rely heavily on knowledge distillation from powerful but inefficient large LMs (LLMs). In this work, we explore a new direction that avoids over-reliance on LLM teachers, introducing a multi-view fine-tuning method that efficiently exploits existing mathematical problem datasets with diverse annotation styles. Our approach uniquely considers the various annotation formats as different "views" and leverages them in training the model. By postpending distinct instructions to input questions, models can learn to generate solutions in diverse formats in a flexible manner. Experimental results show that our strategy enables a LLaMA-7B model to outperform prior approaches that utilize knowledge distillation, as well as carefully established baselines. Additionally, the proposed method grants the models promising generalization ability across various views and datasets, and the capability to learn from inaccurate or incomplete noisy data. We hope our multi-view training paradigm could inspire future studies in other machine reasoning domains.
翻译:数学领域的推理对于相对较小的语言模型(LMs)仍是一个重大挑战。当前许多方法专注于让语言模型专门化数学推理能力,并严重依赖从强大但效率较低的大型语言模型(LLMs)中进行知识蒸馏。本研究探索了一条避免过度依赖大模型教师的新方向,提出了一种多视图微调方法,能高效利用具有多样注释风格的现有数学问题数据集。该方法独特地将不同的注释格式视为不同的"视图",并在模型训练中加以利用。通过在输入问题后添加不同的指令,模型能够灵活地学习生成多种格式的解答。实验结果表明,该策略使LLaMA-7B模型超越了先前利用知识蒸馏的方法以及精心建立的基准方法。此外,所提方法赋予模型跨不同视图和数据集的出色泛化能力,并能从不准确或不完整的噪声数据中学习。我们希望这种多视图训练范式能够启发未来其他机器推理领域的研究。