Recent advancements in Reinforcement Learning with Verifiable Rewards (RLVR) have significantly improved Large Language Model (LLM) reasoning, yet models often struggle to explore novel trajectories beyond their initial latent space. While offline teacher guidance and entropy-driven strategies have been proposed to address this, they often lack deep integration or are constrained by the model's inherent capacity. In this paper, we propose OGER, a novel framework that unifies offline teacher guidance and online reinforcement learning through a specialized reward modeling lens. OGER employs multi-teacher collaborative training and constructs an auxiliary exploration reward that leverages both offline trajectories and the model's own entropy to incentivize autonomous exploration. Extensive experiments across mathematical and general reasoning benchmarks demonstrate that OGER significantly outperforms competitive baselines, achieving substantial gains in mathematical reasoning while maintaining robust generalization to out-of-domain tasks. We provide a comprehensive analysis of training dynamics and conduct detailed ablation studies to validate the effectiveness of our entropy-aware reward modulation. Our code is available at https://github.com/ecoli-hit/OGER.git.
翻译:近年来,基于可验证奖励的强化学习(RLVR)在大语言模型(LLM)推理方面取得了显著进展,但模型仍难以探索超出其初始潜在空间的新颖轨迹。尽管已有离线教师指导和熵驱动策略被提出以解决此问题,但这些方法通常缺乏深度整合,或受限于模型自身的能力。本文提出OGER这一新颖框架,通过专门的奖励建模视角,将离线教师指导与在线强化学习统一起来。OGER采用多教师协同训练,并构建一个辅助探索奖励,该奖励同时利用离线轨迹和模型自身的熵来激励自主探索。在数学推理与通用推理基准上的大量实验表明,OGER显著优于各类强基线,在数学推理中取得可观增益,同时保持对域外任务的鲁棒泛化能力。我们对训练动态进行了全面分析,并通过详细消融实验验证了我们基于熵感知的奖励调制机制的有效性。我们的代码已开源至https://github.com/ecoli-hit/OGER.git。