We propose Process-Aware Policy Optimization (PAPO), a method that integrates process-level evaluation into Group Relative Policy Optimization (GRPO) through decoupled advantage normalization, to address two limitations of existing reward designs. Outcome reward models (ORM) evaluate only final-answer correctness, treating all correct responses identically regardless of reasoning quality, and gradually lose the advantage signal as groups become uniformly correct. Process reward models (PRM) offer richer supervision, but directly using PRM scores causes reward hacking, where models exploit verbosity to inflate scores while accuracy collapses. PAPO resolves both by composing the advantage from an outcome component Aout, derived from ORM and normalized over all responses, and a process component Aproc, derived from a rubric-based PRM and normalized exclusively among correct responses. This decoupled design ensures that Aout anchors training on correctness while Aproc differentiates reasoning quality without distorting the outcome signal. Experiments across multiple model scales and six benchmarks demonstrate that PAPO consistently outperforms ORM, reaching 51.3% vs.\ 46.3% on OlympiadBench while continuing to improve as ORM plateaus and declines.


翻译:我们提出过程感知策略优化(PAPO),该方法通过解耦优势归一化将过程级评估融入组相对策略优化(GRPO),以解决现有奖励设计的两个局限性。结果奖励模型(ORM)仅评估最终答案的正确性,对所有正确响应一视同仁而不考虑推理质量,并且当组内响应逐渐趋于一致正确时会丧失优势信号。过程奖励模型(PRM)提供更丰富的监督信号,但直接使用PRM分数会导致奖励攻击,即模型利用冗长性来抬高分数而准确率却崩溃。PAPO通过将优势分解为两部分来解决这两个问题:结果分量Aout(基于ORM推导并在所有响应上归一化)和过程分量Aproc(基于评分细则的PRM推导并仅在正确响应内部归一化)。这种解耦设计确保Aout锚定正确性训练,而Aproc在区分推理质量的同时不扭曲结果信号。跨多个模型规模和六个基准的实验表明,PAPO一致优于ORM,在OlympiadBench上达到51.3%对46.3%,并且在ORM趋于平稳和下降时仍持续提升。

0
下载
关闭预览

相关内容

《可解释性强化学习模型》
专知会员服务
26+阅读 · 2月24日
【NeurIPS2025】熵正则化与分布强化学习的收敛定理
专知会员服务
12+阅读 · 2025年10月12日
多样化偏好优化
专知会员服务
12+阅读 · 2025年2月3日
直接偏好优化中的数据集、理论、变体和应用的综合综述
专知会员服务
15+阅读 · 2024年10月24日
基于模型的强化学习综述
专知会员服务
48+阅读 · 2023年1月9日
基于模型的强化学习综述
专知
42+阅读 · 2022年7月13日
详解GAN的谱归一化(Spectral Normalization)
PaperWeekly
11+阅读 · 2019年2月13日
一文了解强化学习
AI100
15+阅读 · 2018年8月20日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
24+阅读 · 2015年12月31日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
Arxiv
0+阅读 · 2月20日
VIP会员
相关主题
最新内容
《跨域指挥背景下的领导力发展》最新报告
专知会员服务
0+阅读 · 8分钟前
俄乌无人机战争的六大启示
专知会员服务
10+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
8+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
《履带式无人地面战车技术发展现状》
专知会员服务
8+阅读 · 8月2日
相关基金
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
24+阅读 · 2015年12月31日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员