We propose a novel definition of model exploitation in reinforcement learning. Informally, a world model is exploitable if it implies that one policy should be strictly preferred over another while the environment's true transition model implies the reverse. We analogize our definition with a prior characterization of reward hacking but show that the associated proof of inevitability does not transfer to exploitation. To overcome this obstruction, we develop a general theory of reward hacking and model exploitation that proves that exploitation is essentially unavoidable on large policy sets and yields the corresponding claim for hacking as a special case. Unfortunately, we also find that the conditions that guarantee unhackability in finite policy sets have no counterpart that precludes exploitation. Consequently, we introduce a relaxed notion of exploitation and derive a safe horizon within which it can be avoided. Taken together, our results establish a formal bridge between reward hacking and model exploitation and elucidate the limits of safe planning in world models.
翻译:我们提出了强化学习中模型利用问题的新定义。通俗而言,若世界模型隐含地表明某一策略应严格优于另一策略,而环境的真实转移模型却给出相反的结论,则该世界模型是可被利用的。我们将该定义与奖励欺骗的既有刻画进行类比,但证明其必然性结论无法直接推广至模型利用。为克服这一障碍,我们发展了奖励欺骗与模型利用的统一理论,证明在大型策略集上模型利用本质上不可避免,并推导出欺骗情形作为特例对应的结论。遗憾的是,我们还发现有限策略集中保证不可伪造性的条件不存在对应的可排除模型利用的约束。据此,我们引入松弛化的模型利用概念,并推导出可规避该问题的安全规划时域。综合来看,我们的研究建立了奖励欺骗与模型利用之间的正式联系,阐明了基于世界模型进行安全规划的理论边界。