A growing line of work reframes preference-based fine-tuning of large language models game-theoretically: Nash Learning from Human Feedback (NLHF) recasts the problem as a zero-sum game over policies. However, optimization is over expected pairwise payoffs, thereby conflating policies with similar win rates but different tail behavior. As such, these methods are agnostic to where in the data distribution they succeed or fail: strong average performance can mask systematic failure across prompts, annotators, or safety-critical strata. We introduce risk-sensitive preference games, in which players optimize convex risk measures of their preference loss, exploiting structure in preference uncertainty. While risk-sensitivity generally breaks the zero-sum structure, we show that translation invariance of many risk metrics ensures that we retain monotonicity, yielding fast convergence of sample-efficient self-play methods. Furthermore, we establish algorithmic stability and offline sample complexity bounds that scale with risk, requiring simultaneous control of structural bias from nonlinear risk transformations, statistical bias in risk estimation, and concentration tailored to the risk-sensitive setting. To address statistical bias, we introduce a hierarchical game formulation and a two-timescale extragradient algorithm with bias correction that converges to the Stackelberg equilibrium and is especially effective in low-sample regimes. Empirically, risk-adjusted policies are robust across data strata, stable across risk choices, and match or exceed risk-neutral performance thereby achieving robustness without a performance tax.
翻译:越来越多的研究从博弈论角度重新审视大语言模型的偏好微调:基于人类反馈的纳什学习(NLHF)将这一问题重新表述为策略上的零和博弈。然而,由于优化工作聚焦于期望的成对收益,该方法会将具有相似胜率但尾部行为不同的策略混为一谈。因此,这些方法对其在数据分布中成功或失败的位置不敏感:强大的平均性能可能掩盖在提示、标注者或安全关键层级上的系统性失败。我们引入了风险敏感偏好博弈,其中玩家优化其偏好损失的凸风险度量,从而利用偏好不确定性中的结构。虽然风险敏感性通常会打破零和结构,但我们证明许多风险度量的平移不变性确保了单调性的保留,从而使得样本高效的自对弈方法能够快速收敛。此外,我们建立了与风险成比例缩放的算法稳定性和离线样本复杂度界限,这需要同时控制来自非线性风险变换的结构偏差、风险估计中的统计偏差,以及针对风险敏感环境定制的集中性。为解决统计偏差,我们引入了一种层次化博弈公式和一种具有偏差校正的双时间尺度外梯度算法,该算法收敛至斯塔克尔伯格均衡,且在低样本场景下尤为有效。实验上,风险调整后的策略在各个数据层上具有鲁棒性,跨风险选择保持稳定,并且达到或超过风险中性的性能,从而在无性能代价的前提下实现鲁棒性。