A crucial task in decision-making problems is reward engineering. It is common in practice that no obvious choice of reward function exists. Thus, a popular approach is to introduce human feedback during training and leverage such feedback to learn a reward function. Among all policy learning methods that use human feedback, preference-based methods have demonstrated substantial success in recent empirical applications such as InstructGPT. In this work, we develop a theory that provably shows the benefits of preference-based methods in offline contextual bandits. In particular, we improve the modeling and suboptimality analysis for running policy learning methods on human-scored samples directly. Then, we compare it with the suboptimality guarantees of preference-based methods and show that preference-based methods enjoy lower suboptimality.
翻译:决策问题中的一项关键任务是奖励工程。在实践中,通常不存在显而易见的奖励函数选择。因此,一种常见做法是在训练过程中引入人类反馈,并利用此类反馈来学习奖励函数。在所有使用人类反馈的策略学习方法中,基于偏好的方法在最近的实证应用(如InstructGPT)中已展现出显著成功。本研究发展了一套理论,可证明地展示了基于偏好的方法在离线上下文赌博机中的优势。具体而言,我们改进了直接在人类评分样本上运行策略学习方法的建模与次优性分析,随后将其与基于偏好方法的次优性保证进行比较,表明基于偏好方法具有更低的次优性。