Chain-of-thought (CoT) is a method that enables language models to handle complex reasoning tasks by decomposing them into simpler steps. Despite its success, the underlying mechanics of CoT are not yet fully understood. In an attempt to shed light on this, our study investigates the impact of CoT on the ability of transformers to in-context learn a simple to study, yet general family of compositional functions: multi-layer perceptrons (MLPs). In this setting, we reveal that the success of CoT can be attributed to breaking down in-context learning of a compositional function into two distinct phases: focusing on data related to each step of the composition and in-context learning the single-step composition function. Through both experimental and theoretical evidence, we demonstrate how CoT significantly reduces the sample complexity of in-context learning (ICL) and facilitates the learning of complex functions that non-CoT methods struggle with. Furthermore, we illustrate how transformers can transition from vanilla in-context learning to mastering a compositional function with CoT by simply incorporating an additional layer that performs the necessary filtering for CoT via the attention mechanism. In addition to these test-time benefits, we highlight how CoT accelerates pretraining by learning shortcuts to represent complex functions and how filtering plays an important role in pretraining. These findings collectively provide insights into the mechanics of CoT, inviting further investigation of its role in complex reasoning tasks.
翻译:思维链(Chain-of-Thought,CoT)是一种通过将复杂推理任务分解为更简单步骤来提升语言模型处理能力的方法。尽管该方法已取得显著成功,但其内在机制尚未被完全理解。为阐明这一问题,本研究探讨了CoT对Transformer在上下文学习中处理一类易于研究但具有通用性的组合函数——多层感知机(MLPs)——的影响。在此设定下,我们揭示CoT的成功可归因于将组合函数的上下文学习分解为两个不同阶段:聚焦于组合中每一步骤的相关数据,以及对单步组合函数的上下文学习。通过实验与理论证据,我们展示了CoT如何显著降低上下文学习(ICL)的样本复杂度,并促进非CoT方法难以掌握的复杂函数的学习。此外,我们阐释了Transformer如何通过简单引入一个利用注意力机制执行必要过滤的额外层,从普通上下文学习过渡到使用CoT掌握组合函数。除测试阶段的优势外,我们还着重指出了CoT如何通过学习表征复杂函数的捷径来加速预训练过程,以及过滤机制在预训练中的重要作用。这些发现共同为理解CoT机理提供了洞见,并邀请学界进一步探究其在复杂推理任务中的作用。