Past analyses of reinforcement learning from human feedback (RLHF) assume that the human fully observes the environment. What happens when human feedback is based only on partial observations? We formally define two failure cases: deception and overjustification. Modeling the human as Boltzmann-rational w.r.t. a belief over trajectories, we prove conditions under which RLHF is guaranteed to result in policies that deceptively inflate their performance, overjustify their behavior to make an impression, or both. To help address these issues, we mathematically characterize how partial observability of the environment translates into (lack of) ambiguity in the learned return function. In some cases, accounting for partial observability makes it theoretically possible to recover the return function and thus the optimal policy, while in other cases, there is irreducible ambiguity. We caution against blindly applying RLHF in partially observable settings and propose research directions to help tackle these challenges.
翻译:过去对基于人类反馈的强化学习(RLHF)分析假设人类能完全观测环境。当人类反馈仅基于部分观测时会发生什么?我们正式定义了两种失败案例:欺骗和过度合理化。将人类建模为对轨迹信念具有玻尔兹曼理性后,我们证明了在某些条件下RLHF必然导致策略出现欺骗性夸大性能、为留下印象而过度合理化其行为,或同时出现这两种问题。为帮助解决这些问题,我们从数学上刻画了环境部分可观测性如何转化为学习回报函数中的(不)确定性。在某些情况下,考虑部分可观测性使理论上恢复回报函数及最优策略成为可能,而在其他情况下则存在不可消除的模糊性。我们警示在部分可观测设置中盲目应用RLHF,并提出研究方向以应对这些挑战。