Despite remarkable performance on complex tasks, Large Reasoning Models (LRMs) often generate excessively long Chain-of-Thoughts (CoT), inflating computational costs even for simple queries. Existing efforts to mitigate this inefficiency typically rely on discrete reasoning modes or fixed budget tiers, lacking a principled criterion of when reasoning is sufficient. In this work, we introduce Minimal Sufficient CoT (MSC), defined as the shortest prefix of a CoT trajectory which is adequate for producing the correct answer. We empirically show that MSC not only reduces reasoning tokens, but also improves accuracy across difficulty levels. Building on MSC, we propose Sufficiency-guided Continuous Adaptive Reasoning (SuCo), a two-stage training framework for autonomous reasoning control along a continuous spectrum. In stage 1, MSC-Aligned Fine-Tuning (MFT) constructs MSC data using problem-adaptive sufficiency thresholds that naturally scale with question difficulty, then fine-tunes the model to internalize concise yet sufficient reasoning patterns. In stage 2, Sufficiency-Aware Policy Optimization (SAPO) further optimizes the model through reinforcement learning with dynamic complexity tracking and sufficiency-aware rewards that penalize both over- and under-thinking. Extensive experiments across mathematics, code, and science benchmarks show that SuCo consistently achieves improvements in both accuracy and reasoning efficiency.
翻译:摘要:尽管大型推理模型在复杂任务上表现出色,但其往往生成过长的思维链,即便是简单查询也会显著增加计算成本。现有缓解这一低效问题的方法通常依赖于离散推理模式或固定预算层级,缺乏推理何时足够充分的原理性判据。本文提出最小充分思维链(MSC),定义为思维链轨迹中足以产生正确答案的最短前缀。实验表明,MSC不仅能减少推理令牌数量,还能跨难度级别提升准确率。基于MSC,我们提出充分性引导的连续自适应推理(SuCo),一种面向连续谱系实现推理自主控制的两阶段训练框架。第一阶段:MSC对齐微调(MFT)利用与问题难度自然关联的自适应充分性阈值构建MSC数据,然后微调模型使其内化简洁而充分的推理模式。第二阶段:充分性感知策略优化(SAPO)通过结合动态复杂度追踪和充分性感知奖励(同时惩罚过度思考与思考不足)的强化学习进一步优化模型。在数学、代码和科学基准上的大量实验表明,SuCo在准确率和推理效率两方面均持续实现改进。