Large Language Models (LLMs) have revolutionized natural language processing, yet aligning these models with human values and preferences using RLHF remains a significant challenge. This challenge is characterized by various instabilities, such as reward hacking and catastrophic forgetting. In this technical report, we propose two innovations to stabilize RLHF training: 1) Advantage Model, which directly models advantage score i.e., extra reward compared to the expected rewards and regulates score distributions across tasks to prevent reward hacking. 2) Selective Rehearsal, which mitigates catastrophic forgetting by strategically selecting data for PPO training and knowledge rehearsing. Our experimental analysis on public and proprietary datasets reveals that the proposed methods not only increase stability in RLHF training but also achieve higher reward scores and win rates.
翻译:大型语言模型(LLMs)革新了自然语言处理领域,但利用RLHF将这些模型与人类价值观和偏好对齐仍是一项重大挑战。该挑战表现为多种不稳定性,例如奖励欺骗和灾难性遗忘。在本技术报告中,我们提出两项创新以稳定RLHF训练:1)优势模型(Advantage Model),该模型直接建模优势分数(即相对于期望奖励的额外奖励),并调节任务间的分数分布以防止奖励欺骗。2)选择性重演(Selective Rehearsal),通过策略性地选择PPO训练和知识重演的数据来缓解灾难性遗忘。我们在公开和专有数据集上的实验分析表明,所提出的方法不仅提高了RLHF训练的稳定性,还获得了更高的奖励分数和胜率。