Large language models (LLMs) can achieve highly effective performance on various reasoning tasks by incorporating step-by-step chain-of-thought (CoT) prompting as demonstrations. However, the reasoning chains of demonstrations generated by LLMs are prone to errors, which can subsequently lead to incorrect reasoning during inference. Furthermore, inappropriate exemplars (overly simplistic or complex), can affect overall performance among varying levels of difficulty. We introduce Iter-CoT (Iterative bootstrapping in Chain-of-Thoughts Prompting), an iterative bootstrapping approach for selecting exemplars and generating reasoning chains. By utilizing iterative bootstrapping, our approach enables LLMs to autonomously rectify errors, resulting in more precise and comprehensive reasoning chains. Simultaneously, our approach selects challenging yet answerable questions accompanied by reasoning chains as exemplars with a moderate level of difficulty, which enhances the LLMs' generalizability across varying levels of difficulty. Experimental results indicate that Iter-CoT exhibits superiority, achieving competitive performance across three distinct reasoning tasks on eleven datasets.
翻译:摘要:大语言模型(LLMs)通过融入逐步的思维链(Chain-of-Thought, CoT)提示作为示例,可在各类推理任务中实现高效性能。然而,由LLMs生成的示例推理链容易出错,进而导致推理过程中的错误推导。此外,不恰当的示例(过于简单或复杂)可能影响模型在不同难度层级上的整体表现。我们提出Iter-CoT(迭代自举思维链提示),一种通过迭代自举选取示例并生成推理链的方法。通过运用迭代自举,该方法使LLMs能够自主纠正错误,生成更精确且全面的推理链。同时,我们的方法选取具有中等难度、兼具挑战性且可回答的问题及其推理链作为示例,从而提升LLMs在不同难度层级上的泛化能力。实验结果表明,Iter-CoT展现出优越性,在三个不同推理任务的十一个数据集上均取得了具有竞争力的性能。