Existing algorithms for reinforcement learning from human feedback (RLHF) can incentivize responses at odds with preferences because they are based on models that assume independence of irrelevant alternatives (IIA). The perverse incentives induced by IIA give rise to egregious behavior when innovating on query formats or learning algorithms.
翻译:现有基于人类反馈的强化学习(RLHF)算法可能激励与偏好相悖的响应,其根源在于这些算法所依赖的模型假设了无关选项独立性(IIA)。IIA引发的反常激励在查询格式创新或学习算法改进时会导致严重异常行为。