For sparse, structured reinforcement-learning tasks with semantic reward-function interfaces, LLM-generated reward shaping is better framed as debugging than one-shot generation. We study PPO-trained agents using MiniGrid as core evaluation and MuJoCo as boundary stress test. Our audit finds two dominant one-shot failure modes -- reward flooding and semantic/API misunderstanding -- plus a rarer weak-shaping case. We propose diagnostic-driven iterative refinement, where training diagnostics and a failure-mode taxonomy guide targeted reward-function revision. Refinement improves DoorKey-8x8 from 2.3% to 97.6% and KeyCorridor from 31.2% to 86.7% with high seed-to-seed variance. Controls show these gains are not from retrying or extra training: metrics-only re-prompting yields large drops, while a static-vocabulary control recovers much of the gap (87.6%; 70.7%), showing the taxonomy prompt is a major mechanism and dynamic labels provide only partially isolated incremental evidence. Budget-matched and Best-of-3 comparisons separate refinement from selection and training-time effects. Component-removal tests, sensitivity analyses, and an audit against author labels provide converging evidence for the debugging interpretation while revealing calibration limits. Continuous-control results show the boundary: success-based diagnostics can misfire in dense-reward locomotion, and return-trend feedback removes one false-positive mechanism without robust gains. The low-call protocol is a cost contrast with population-based reward search, not a benchmark comparison. In four crossed-variance-design environments, point estimates suggest larger gains when LLM reward-function variance dominates but bootstrap intervals are wide. The method is bounded to sparse structured tasks with reliable interfaces under PPO; fields like event_text may help, hurt, or be neutral.


翻译:对于具有语义奖励函数接口的稀疏、结构化强化学习任务而言,将大型语言模型生成的奖励塑造视为一种调试过程,而非一次性生成任务更为合理。我们以MiniGrid作为核心评估环境、MuJoCo作为边界压力测试,研究了采用PPO算法训练的智能体。审计发现两种主导性的一次性生成失败模式——奖励泛滥与语义/应用程序编程接口误解(API misunderstanding),以及一种罕见的弱塑造案例。我们提出基于诊断驱动的迭代优化方法,利用训练诊断结果与失败模式分类法指导针对性奖励函数修订。在DoorKey-8x8环境中,该方法将成功率从2.3%提升至97.6%,在KeyCorridor环境中从31.2%提升至86.7%,但种子间方差较大。对照实验表明,这些提升并非源于重试或额外训练:仅基于指标的重新提示会导致大幅性能下降,而静态词汇控制方法可恢复大部分差距(87.6%;70.7%),这证实分类提示是主要机制,而动态标签仅提供部分孤立的增量证据。预算匹配对比与最佳取三(Best-of-3)比较将优化方法与选择效应和训练时间效应相分离。组件消融测试、敏感性分析以及与作者标签的核对审计为调试解释提供了收敛性证据,同时揭示了校准局限性。连续控制结果展示了边界情况:基于成功率的诊断在密集奖励运动控制任务中可能失效,而回报趋势反馈虽消除了一个假阳性机制,但未带来稳健的性能提升。低调用协议是与基于种群的奖励搜索进行成本对比,而非基准性能比较。在四类交叉方差设计环境中,点估计表明当LLM奖励函数方差占主导时收益更大,但自助法置信区间较宽。该方法仅限于在PPO算法下具有可靠接口的稀疏结构化任务;诸如事件文本等字段可能产生正向、负向或中性影响。

0
下载
关闭预览

相关内容

【EMNLP2025】面向大语言模型的权重旋转偏好优化
专知会员服务
12+阅读 · 2025年8月27日
深度强化学习中的奖励模型:综述
专知会员服务
29+阅读 · 2025年6月20日
【博士论文】强化学习智能体的奖励函数设计
专知会员服务
49+阅读 · 2025年4月8日
Llama-3-SynE:实现有效且高效的大语言模型持续预训练
专知会员服务
36+阅读 · 2024年7月30日
基于模型的强化学习综述
专知
42+阅读 · 2022年7月13日
强化学习《奖励函数设计: Reward Shaping》详细解读
深度强化学习实验室
20+阅读 · 2020年9月1日
Distributional Soft Actor-Critic (DSAC)强化学习算法的设计与验证
深度强化学习实验室
20+阅读 · 2020年8月11日
以BERT为例,如何优化机器学习模型性能?
专知
10+阅读 · 2019年10月3日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
Arxiv
0+阅读 · 6月15日
VIP会员
最新内容
博士论文 | 用代码结构感知方法推进代码大模型
《决策模型比较研究》
专知会员服务
8+阅读 · 7月25日
《美军水下战与海床战概述及本地实施》
专知会员服务
6+阅读 · 7月25日
面向未来冲突推进陆军情报体制改革
专知会员服务
4+阅读 · 7月25日
乌克兰纵深打击如何重塑俄罗斯的战略选择
专知会员服务
3+阅读 · 7月24日
俄乌战争中关于中程打击无人机部署的经验启示
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员