Language models can be prompted to reason through problems in a manner that significantly improves performance. However, \textit{why} such prompting improves performance is unclear. Recent work showed that using logically \textit{invalid} Chain-of-Thought (CoT) prompting improves performance almost as much as logically \textit{valid} CoT prompting, and that editing CoT prompts to replace problem-specific information with abstract information or out-of-distribution information typically doesn't harm performance. Critics have responded that these findings are based on too few and too easy tasks to draw meaningful conclusions. To resolve this dispute, we test whether logically invalid CoT prompts offer the same level of performance gains as logically valid prompts on the hardest tasks in the BIG-Bench benchmark, termed BIG-Bench Hard (BBH). We find that the logically \textit{invalid} reasoning prompts do indeed achieve similar performance gains on BBH tasks as logically valid reasoning prompts. We also discover that some CoT prompts used by previous works contain logical errors. This suggests that covariates beyond logically valid reasoning are responsible for performance improvements.
翻译:语言模型可以通过提示以推理方式处理问题,从而显著提升性能。然而,为何此类提示能改善性能尚不明确。近期研究表明,使用逻辑上无效的思维链提示几乎与逻辑上有效的思维链提示一样能提升性能,且编辑思维链提示、将问题特定信息替换为抽象信息或分布外信息通常不会损害性能。批评者回应称,这些发现基于过少且过于简单的任务,无法得出有意义的结论。为解决这一争议,我们测试了在BIG-Bench基准中最困难任务(称为BIG-Bench Hard,BBH)上,逻辑无效的思维链提示是否与逻辑有效提示提供同等水平的性能增益。我们发现,逻辑无效的推理提示确实在BBH任务上实现了与逻辑有效推理提示相似的性能增益。我们还发现,先前工作中的某些思维链提示存在逻辑错误。这表明,除逻辑有效推理之外的协变量才是性能提升的原因。