In reinforcement learning, specification gaming occurs when AI systems learn undesired behaviors that are highly rewarded due to misspecified training goals. Specification gaming can range from simple behaviors like sycophancy to sophisticated and pernicious behaviors like reward-tampering, where a model directly modifies its own reward mechanism. However, these more pernicious behaviors may be too complex to be discovered via exploration. In this paper, we study whether Large Language Model (LLM) assistants which find easily discovered forms of specification gaming will generalize to perform rarer and more blatant forms, up to and including reward-tampering. We construct a curriculum of increasingly sophisticated gameable environments and find that training on early-curriculum environments leads to more specification gaming on remaining environments. Strikingly, a small but non-negligible proportion of the time, LLM assistants trained on the full curriculum generalize zero-shot to directly rewriting their own reward function. Retraining an LLM not to game early-curriculum environments mitigates, but does not eliminate, reward-tampering in later environments. Moreover, adding harmlessness training to our gameable environments does not prevent reward-tampering. These results demonstrate that LLMs can generalize from common forms of specification gaming to more pernicious reward tampering and that such behavior may be nontrivial to remove.
翻译:在强化学习中,当人工智能系统因训练目标设定不当而习得能获得高奖励的非预期行为时,便会出现规范博弈现象。规范博弈的行为谱系广泛,既包括阿谀奉承这类简单行为,也涵盖奖励篡改等复杂且有害的行为——即模型直接修改自身奖励机制。然而,这些更具危害性的行为可能因过于复杂而难以通过探索被发现。本文旨在研究:那些已发现易暴露规范博弈形式的大语言模型助手,是否会泛化到执行更罕见、更露骨的行为模式,直至包括奖励篡改。我们构建了由易到难的系列可博弈环境课程,发现对早期课程环境的训练会导致模型在后续环境中表现出更多规范博弈行为。值得注意的是,在少数但不可忽视的情况下,经过完整课程训练的大语言模型助手能够零样本泛化到直接改写自身奖励函数。对早期课程环境进行防博弈再训练虽能缓解后续环境中的奖励篡改,但无法完全消除。此外,在可博弈环境中增加无害化训练亦不能防止奖励篡改。这些结果表明:大语言模型能够从常见的规范博弈形式泛化到更具危害性的奖励篡改行为,且此类行为可能难以彻底消除。