For sparse, structured reinforcement-learning tasks with semantic reward-function interfaces, LLM-generated reward shaping is better framed as debugging than one-shot generation. We study PPO-trained agents using MiniGrid as core evaluation and MuJoCo as boundary stress test. Our audit finds two dominant one-shot failure modes -- reward flooding and semantic/API misunderstanding -- plus a rarer weak-shaping case. We propose diagnostic-driven iterative refinement, where training diagnostics and a failure-mode taxonomy guide targeted reward-function revision. Refinement improves DoorKey-8x8 from 2.3% to 97.6% and KeyCorridor from 31.2% to 86.7% with high seed-to-seed variance. Controls show these gains are not from retrying or extra training: metrics-only re-prompting yields large drops, while a static-vocabulary control recovers much of the gap (87.6%; 70.7%), showing the taxonomy prompt is a major mechanism and dynamic labels provide only partially isolated incremental evidence. Budget-matched and Best-of-3 comparisons separate refinement from selection and training-time effects. Component-removal tests, sensitivity analyses, and an audit against author labels provide converging evidence for the debugging interpretation while revealing calibration limits. Continuous-control results show the boundary: success-based diagnostics can misfire in dense-reward locomotion, and return-trend feedback removes one false-positive mechanism without robust gains. The low-call protocol is a cost contrast with population-based reward search, not a benchmark comparison. In four crossed-variance-design environments, point estimates suggest larger gains when LLM reward-function variance dominates but bootstrap intervals are wide. The method is bounded to sparse structured tasks with reliable interfaces under PPO; fields like event_text may help, hurt, or be neutral.
翻译:对于具有语义奖励函数接口的稀疏、结构化强化学习任务而言,将大型语言模型生成的奖励塑造视为一种调试过程,而非一次性生成任务更为合理。我们以MiniGrid作为核心评估环境、MuJoCo作为边界压力测试,研究了采用PPO算法训练的智能体。审计发现两种主导性的一次性生成失败模式——奖励泛滥与语义/应用程序编程接口误解(API misunderstanding),以及一种罕见的弱塑造案例。我们提出基于诊断驱动的迭代优化方法,利用训练诊断结果与失败模式分类法指导针对性奖励函数修订。在DoorKey-8x8环境中,该方法将成功率从2.3%提升至97.6%,在KeyCorridor环境中从31.2%提升至86.7%,但种子间方差较大。对照实验表明,这些提升并非源于重试或额外训练:仅基于指标的重新提示会导致大幅性能下降,而静态词汇控制方法可恢复大部分差距(87.6%;70.7%),这证实分类提示是主要机制,而动态标签仅提供部分孤立的增量证据。预算匹配对比与最佳取三(Best-of-3)比较将优化方法与选择效应和训练时间效应相分离。组件消融测试、敏感性分析以及与作者标签的核对审计为调试解释提供了收敛性证据,同时揭示了校准局限性。连续控制结果展示了边界情况:基于成功率的诊断在密集奖励运动控制任务中可能失效,而回报趋势反馈虽消除了一个假阳性机制,但未带来稳健的性能提升。低调用协议是与基于种群的奖励搜索进行成本对比,而非基准性能比较。在四类交叉方差设计环境中,点估计表明当LLM奖励函数方差占主导时收益更大,但自助法置信区间较宽。该方法仅限于在PPO算法下具有可靠接口的稀疏结构化任务;诸如事件文本等字段可能产生正向、负向或中性影响。