Let $p$ denote a generative language model. Let $r$ denote a reward model that returns a scalar that captures the degree at which a draw from $p$ is preferred. The goal of language model alignment is to alter $p$ to a new distribution $\phi$ that results in a higher expected reward while keeping $\phi$ close to $p.$ A popular alignment method is the KL-constrained reinforcement learning (RL), which chooses a distribution $\phi_\Delta$ that maximizes $E_{\phi_{\Delta}} r(y)$ subject to a relative entropy constraint $KL(\phi_\Delta || p) \leq \Delta.$ Another simple alignment method is best-of-$N$, where $N$ samples are drawn from $p$ and one with highest reward is selected. In this paper, we offer a closed-form characterization of the optimal KL-constrained RL solution. We demonstrate that any alignment method that achieves a comparable trade-off between KL divergence and reward must approximate the optimal KL-constrained RL solution in terms of relative entropy. To further analyze the properties of alignment methods, we introduce two simplifying assumptions: we let the language model be memoryless, and the reward model be linear. Although these assumptions may not reflect complex real-world scenarios, they enable a precise characterization of the asymptotic behavior of both the best-of-$N$ alignment, and the KL-constrained RL method, in terms of information-theoretic quantities. We prove that the reward of the optimal KL-constrained RL solution satisfies a large deviation principle, and we fully characterize its rate function. We also show that the rate of growth of the scaled cumulants of the reward is characterized by a proper Renyi cross entropy. Finally, we show that best-of-$N$ is asymptotically equivalent to KL-constrained RL solution by proving that their expected rewards are asymptotically equal, and concluding that the two distributions must be close in KL divergence.
翻译:设$p$表示生成式语言模型,$r$表示返回标量奖励的奖励模型,该标量衡量从$p$中抽取样本的偏好程度。语言模型对齐的目标是将$p$调整为新的分布$\phi$,在保持$\phi$接近$p$的同时获得更高期望奖励。一种流行的对齐方法为KL约束强化学习(RL),它在满足相对熵约束$KL(\phi_\Delta || p) \leq \Delta$的条件下,选择最大化$E_{\phi_{\Delta}} r(y)$的分布$\phi_\Delta$。另一种简单的对齐方法为最佳-$N$采样,即从$p$中抽取$N$个样本并选择奖励最高的样本。本文给出了最优KL约束RL解的闭式刻画,证明任何能在KL散度与奖励之间实现可比较权衡的对齐方法,其相对熵必然近似于最优KL约束RL解。为深入分析对齐方法性质,我们引入两个简化假设:语言模型为无记忆模型,奖励模型为线性模型。虽然这些假设可能无法反映复杂现实场景,但它们使我们能够从信息论角度精确刻画最佳-$N$对齐与KL约束RL方法的渐近行为。我们证明最优KL约束RL解的奖励满足大偏差原理,并完整刻画其速率函数。同时表明奖励的缩放累积量增长率由恰当的Rényi交叉熵决定。最后,通过证明最佳-$N$与KL约束RL解的期望奖励渐近相等,并得出两个分布必须在KL散度上接近,从而证实最佳-$N$渐近等价于KL约束RL解。