Group-based reinforcement learning methods, like Group Relative Policy Optimization (GRPO), are widely used nowadays to post-train large language models. Despite their empirical success, they exhibit structural mismatches between reward optimization and the underlying training objective. In this paper, we present a theoretical analysis of GRPO style methods by studying them within a unified surrogate formulation. This perspective reveals recurring properties that affect all the methods under analysis: (i) non-uniform group weighting induces systematic gradient biases on shared prefix tokens; (ii) interactions with the AdamW optimizer make training dynamics largely insensitive to reward scaling; and (iii) optimizer momentum can push policy updates beyond the intended clipping region under repeated optimization steps. We believe that these findings highlight fundamental limitations of current approaches and provide principled guidance for the design of future formulations.


翻译:基于群体的强化学习方法,如群体相对策略优化(GRPO),如今被广泛用于大型语言模型的后训练。尽管它们在实证上取得了成功,但这些方法在奖励优化与底层训练目标之间存在着结构性的不匹配。本文通过对GRPO风格方法在一个统一的代理公式框架内进行研究,提出了理论分析。这一视角揭示了影响所有被分析方法的反复出现的特性:(i)非均匀的群体加权会在共享前缀词元上引入系统性的梯度偏差;(ii)与AdamW优化器的交互使得训练动态在很大程度上对奖励缩放不敏感;以及(iii)在重复的优化步骤下,优化器的动量可能将策略更新推至预期的裁剪区域之外。我们相信,这些发现突显了当前方法的根本性局限,并为未来公式的设计提供了原则性指导。

0
下载
关闭预览

相关内容

《深度强化学习在集群系统中的应用》31页论文
专知会员服务
61+阅读 · 2023年3月14日
基于模型的强化学习综述
专知会员服务
150+阅读 · 2022年7月13日
基于模型的强化学习综述
专知
42+阅读 · 2022年7月13日
探索(Exploration)还是利用(Exploitation)?强化学习如何tradeoff?
深度强化学习实验室
13+阅读 · 2020年8月23日
强化学习的两大话题之一,仍有极大探索空间
AI科技评论
22+阅读 · 2020年8月22日
强化学习开篇:Q-Learning原理详解
AINLP
37+阅读 · 2020年7月28日
论强化学习的根本缺陷
AI科技评论
11+阅读 · 2018年7月24日
关于强化学习(附代码,练习和解答)
深度学习
38+阅读 · 2018年1月30日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
41+阅读 · 2015年12月31日
国家自然科学基金
24+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
Arxiv
0+阅读 · 1月13日
VIP会员
最新内容
《无人机蜂群:释放人类-蜂群编队的潜能》
专知会员服务
0+阅读 · 10分钟前
《战略战术化:一项综合性述评》
专知会员服务
0+阅读 · 14分钟前
美陆军-工业界协同推进反无人机系统技术发展
专知会员服务
1+阅读 · 36分钟前
《跨域指挥背景下的领导力发展》最新报告
专知会员服务
0+阅读 · 42分钟前
俄乌无人机战争的六大启示
专知会员服务
10+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
8+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
41+阅读 · 2015年12月31日
国家自然科学基金
24+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员