To align conditional text generation model outputs with desired behaviors, there has been an increasing focus on training the model using reinforcement learning (RL) with reward functions learned from human annotations. Under this framework, we identify three common cases where high rewards are incorrectly assigned to undesirable patterns: noise-induced spurious correlation, naturally occurring spurious correlation, and covariate shift. We show that even though learned metrics achieve high performance on the distribution of the data used to train the reward function, the undesirable patterns may be amplified during RL training of the text generation model. While there has been discussion about reward gaming in the RL or safety community, in this discussion piece, we would like to highlight reward gaming in the natural language generation (NLG) community using concrete conditional text generation examples and discuss potential fixes and areas for future work.
翻译:为使条件文本生成模型的输出符合预期行为,研究者日益关注使用基于人工标注学习到的奖励函数,通过强化学习(RL)训练模型。在此框架下,我们识别出三种常见的高奖励被错误分配给不良模式的案例:噪声诱发的虚假相关、自然存在的虚假相关以及协变量偏移。研究表明,尽管习得的评估指标在用于训练奖励函数的数据分布上表现优异,但在文本生成模型的强化学习训练过程中,这些不良模式可能被放大。尽管强化学习或安全领域已有关于奖励操控的讨论,但在这篇讨论性文章中,我们旨在通过具体的条件文本生成实例,向自然语言生成(NLG)领域强调奖励操控问题,并探讨潜在的改进方案及未来研究方向。