Reinforcement learning substantially improves reasoning in large language models, but it also tends to lengthen chain-of-thought outputs and increase computational cost. Although length-control methods have been proposed, the length-accuracy relationship they induce remains unclear. We train policies with several length-control methods on multiple base models in a controlled setup and find that, across both mathematical reasoning and code generation, accuracy is non-monotonic in output length, peaking at an intermediate value. Mode accuracy, however, continues to improve with length even in settings where sample accuracy plateaus or declines, indicating that the non-monotonic length-accuracy relationship is driven by dispersion around an increasingly correct center.
翻译:强化学习显著提升了大型语言模型的推理能力,但同时也倾向于延长思维链输出并增加计算成本。尽管已有长度控制方法被提出,但由此引发的长度-准确性关系仍不明确。我们在受控环境下对多种基础模型使用多种长度控制方法训练策略,发现无论是在数学推理还是代码生成任务中,准确性随输出长度呈非单调变化,在中间值达到峰值。然而,即使在样本准确率趋于平缓或下降的场景中,模式准确率仍随长度持续提升,这表明非单调的长度-准确率关系是由围绕日益正确的中心产生的分散性所驱动的。