Reinforcement learning from human feedback (RLHF) has been an effective technique for aligning AI systems with human values, with remarkable successes in fine-tuning large-language models recently. Most existing RLHF paradigms make the underlying assumption that human preferences are relatively homogeneous, and can be encoded by a single reward model. In this paper, we focus on addressing the issues due to the inherent heterogeneity in human preferences, as well as their potential strategic behavior in providing feedback. Specifically, we propose two frameworks to address heterogeneous human feedback in principled ways: personalization-based one and aggregation-based one. For the former, we propose two approaches based on representation learning and clustering, respectively, for learning multiple reward models that trades off the bias (due to preference heterogeneity) and variance (due to the use of fewer data for learning each model by personalization). We then establish sample complexity guarantees for both approaches. For the latter, we aim to adhere to the single-model framework, as already deployed in the current RLHF paradigm, by carefully aggregating diverse and truthful preferences from humans. We propose two approaches based on reward and preference aggregation, respectively: the former utilizes both utilitarianism and Leximin approaches to aggregate individual reward models, with sample complexity guarantees; the latter directly aggregates the human feedback in the form of probabilistic opinions. Under the probabilistic-opinion-feedback model, we also develop an approach to handle strategic human labelers who may bias and manipulate the aggregated preferences with untruthful feedback. Based on the ideas in mechanism design, our approach ensures truthful preference reporting, with the induced aggregation rule maximizing social welfare functions.
翻译:基于人类反馈的强化学习(RLHF)已成为使人工智能系统与人类价值观对齐的有效技术,近期在大型语言模型微调中取得了显著成功。现有RLHF范式大多基于一个潜在假设:人类偏好相对同质,且可通过单一奖励模型编码。本文重点解决因人类偏好固有异质性及其提供反馈时潜在策略行为所引发的问题。具体而言,我们提出两个以原则性方法处理异构人类反馈的框架:基于个性化的框架与基于聚合的框架。针对前者,我们分别提出基于表征学习和聚类的两种方法,用于学习多个奖励模型,以权衡偏差(源于偏好异质性)与方差(源于个性化导致每个模型训练数据减少)。随后为两种方法建立了样本复杂度保证。针对后者,我们旨在遵循当前RLHF范式已部署的单模型框架,通过审慎聚合人类多样且真实的偏好来实现。我们分别提出基于奖励聚合和偏好聚合的两种方法:前者运用功利主义与Leximin方法聚合个体奖励模型,并提供样本复杂度保证;后者直接聚合以概率意见形式呈现的人类反馈。在概率意见反馈模型下,我们还开发了一种处理策略性人类标注者的方法——这些标注者可能通过不真实反馈使聚合偏好产生偏差或被操纵。基于机制设计思想,该方法能确保真实偏好报告,且所导出的聚合规则可最大化社会福利函数。