Large Language Models (LLMs) show remarkable capabilities, yet their stochastic next-token prediction creates logical inconsistencies and reward hacking that formal symbolic systems avoid. To bridge this gap, we introduce a formal logic verification-guided framework that dynamically interleaves formal symbolic verification with the natural language generation process, providing real-time feedback to detect and rectify errors as they occur. Distinguished from previous neuro-symbolic methods limited by passive post-hoc validation, our approach actively penalizes intermediate fallacies during the reasoning chain. We operationalize this framework via a novel two-stage training pipeline that synergizes formal logic verification-guided supervised fine-tuning and policy optimization. Extensive evaluation on six benchmarks spanning mathematical, logical, and general reasoning demonstrates that our 7B and 14B models outperform state-of-the-art baselines by average margins of 10.4% and 14.2%, respectively. These results validate that formal verification can serve as a scalable mechanism to significantly push the performance boundaries of advanced LLM reasoning.
翻译:大型语言模型(LLMs)展现出卓越能力,但其随机性下一个词预测机制易产生逻辑不一致与奖励攻击,而传统形式符号系统可规避此类问题。为弥合这一差距,我们提出形式逻辑验证引导框架,通过动态交错形式符号验证与自然语言生成过程,在推理过程中提供实时反馈以检测并修正错误。与受限于被动事后验证的先前神经符号方法不同,本方法在推理链中主动惩罚中间谬误。我们通过创新的两阶段训练范式实现该框架:协同融合形式逻辑验证引导的监督微调与策略优化。在涵盖数学推理、逻辑推理与通用推理的六个基准测试上的广泛评估表明,我们的7B与14B模型分别以平均10.4%与14.2%的绝对优势超越当前最优基线。这些结果验证了形式验证可作为可扩展机制,显著推升先进LLM推理性能边界。