Aligning intelligent agents with human preferences and values is important. This paper examines two popular alignment methods: Direct Preference Optimization (DPO) and Reward-Model-Based Policy Optimization (RMB-PO). A variant of RMB-PO, referred to as RMB-PO+ is also considered. These methods, either explicitly or implicitly, learn a reward model from preference data and differ in the data used for policy optimization to unlock the generalization ability of the reward model. In particular, compared with DPO, RMB-PO additionally uses policy-generated data, and RMB-PO+ further leverages new, preference-free data. We examine the impact of such out-of-preference data. Our study, conducted through controlled and synthetic experiments, demonstrates that DPO performs poorly, whereas RMB-PO+ performs the best. In particular, even when providing the policy model with a good feature representation, we find that policy optimization with adequate out-of-preference data significantly improves performance by harnessing the reward model's generalization capabilities.
翻译:使智能体与人类偏好及价值观对齐至关重要。本文研究了两种流行的对齐方法:直接偏好优化(DPO)和基于奖励模型的策略优化(RMB-PO),同时考虑了RMB-PO的一种变体,即RMB-PO+。这些方法均显式或隐式地从偏好数据中学习奖励模型,但它们在策略优化过程中使用的数据不同,以释放奖励模型的泛化能力。具体而言,与DPO相比,RMB-PO额外使用策略生成的数据,而RMB-PO+则进一步利用新的、无偏好的数据。我们探究了此类偏好外数据的影响。通过受控实验与合成实验,本研究表明DPO表现较差,而RMB-PO+表现最优。特别是在为策略模型提供良好特征表示的情况下,我们发现利用充足的偏好外数据进行策略优化,可通过发挥奖励模型的泛化能力显著提升性能。