Training on large amounts of rationales (i.e., CoT Fine-tuning) is effective at improving the reasoning capabilities of large language models (LLMs). However, acquiring human-authored rationales or augmenting rationales from proprietary models is costly and not scalable. In this paper, we study the problem of whether LLMs could self-improve their reasoning capabilities. To this end, we propose Self-Explore, where the LLM is tasked to explore the first wrong step (i.e., the first pit) within the rationale and use such signals as fine-grained rewards for further improvement. On the GSM8K and MATH test set, Self-Explore achieves 11.57% and 2.89% improvement on average across three LLMs compared to supervised fine-tuning (SFT). Our code is available at https://github.com/hbin0701/Self-Explore.
翻译:大量基于推理链(即,CoT微调)的训练能有效提升大语言模型的推理能力。然而,获取人工撰写的推理链或从专有模型中增强推理链成本高昂且难以扩展。本文研究了语言模型能否自我提升其推理能力的问题。为此,我们提出Self-Explore方法,要求模型在推理链中探索首个错误步骤(即,首个陷阱),并将此类信号作为细粒度奖励用于进一步改进。在GSM8K和MATH测试集上,与监督微调(SFT)相比,Self-Explore在三种语言模型上平均分别实现了11.57%和2.89%的性能提升。我们的代码已开源至https://github.com/hbin0701/Self-Explore。