Large language models show strong reasoning ability, but their internal reasoning process can remain unstable in complex multi-step settings, where early hidden-state errors may propagate to incorrect predictions. We propose ReLAR, a reinforcement-guided latent refinement framework that iteratively updates hidden representations before decoding. ReLAR maintains a compact latent reasoning state and uses learned depth and action controllers to adaptively determine both the number and direction of refinement steps. The controllers are trained with a policy gradient objective based on step-wise likelihood improvement, enabling efficient input-dependent reasoning without explicit chain-of-thought generation. Experiments on medical, mathematical, multi-hop reasoning, and open-ended generation benchmarks show that ReLAR improves accuracy, generation quality, and reasoning stability with substantially lower inference overhead than explicit reasoning baselines.
翻译:大语言模型展现出强大的推理能力,但在复杂的多步推理场景中,其内部推理过程仍可能不稳定——早期隐藏状态中的错误可能会传播至后续预测。我们提出ReLAR,一种基于强化学习的潜在状态精炼框架,可在解码前迭代更新隐藏表示。ReLAR维护一个紧凑的潜在推理状态,并通过学习得到的深度控制器和动作控制器自适应地确定精炼步骤的数量与方向。控制器基于逐步骤似然提升策略梯度目标进行训练,从而无需显式思维链生成即可实现高效的输入依赖推理。在医学、数学、多步推理及开放式生成任务上的实验表明,与显式推理基线相比,ReLAR在显著降低推理开销的同时提升了准确性、生成质量与推理稳定性。