CoT (Chain-of-Thought) is a way to solve reasoning problems for LLMs . Recently, many researches appear for improving the CoT capability of LLMs. In this work, we also proposed Olapa-MCoT, which is a LLMs based on llama2-13B PLM for finetuning and alignment learning. During the alignment training, we proposed the SimRRHF algorithm and Incorrect Data Relearning and mainly focused on optimizing the Chinese mathematical reasoning ability of Olapa-MCoT. The experiment achieved significant results, with the accuracy of Chinese mathematical reasoning up to 50%, 36% rise compared to llama2-13B. In addition, the accuracy of English reasoning ability also increased by nearly 4%.
翻译:思维链(Chain-of-Thought, CoT)是解决大语言模型推理问题的一种方法。近期,关于提升大语言模型CoT能力的研究层出不穷。本文提出Olapa-MCoT,该模型基于llama2-13B预训练语言模型进行微调与对齐学习。在对齐训练过程中,我们提出了SimRRHF算法与错误数据重学习方法,重点优化Olapa-MCoT的中文数学推理能力。实验取得了显著成果:中文数学推理准确率最高达50%,较llama2-13B提升36%;此外,英文推理能力准确率亦提升近4%。