Generating complex behaviors that satisfy the preferences of non-expert users is a crucial requirement for AI agents. Interactive reward learning from trajectory comparisons (a.k.a. RLHF) is one way to allow non-expert users to convey complex objectives by expressing preferences over short clips of agent behaviors. Even though this parametric method can encode complex tacit knowledge present in the underlying tasks, it implicitly assumes that the human is unable to provide richer feedback than binary preference labels, leading to intolerably high feedback complexity and poor user experience. While providing a detailed symbolic closed-form specification of the objectives might be tempting, it is not always feasible even for an expert user. However, in most cases, humans are aware of how the agent should change its behavior along meaningful axes to fulfill their underlying purpose, even if they are not able to fully specify task objectives symbolically. Using this as motivation, we introduce the notion of Relative Behavioral Attributes, which allows the users to tweak the agent behavior through symbolic concepts (e.g., increasing the softness or speed of agents' movement). We propose two practical methods that can learn to model any kind of behavioral attributes from ordered behavior clips. We demonstrate the effectiveness of our methods on four tasks with nine different behavioral attributes, showing that once the attributes are learned, end users can produce desirable agent behaviors relatively effortlessly, by providing feedback just around ten times. This is over an order of magnitude less than that required by the popular learning-from-human-preferences baselines. The supplementary video and source code are available at: https://guansuns.github.io/pages/rba.
翻译:生成能够满足非专家用户偏好的复杂行为是AI智能体的关键需求。基于轨迹比较的交互式奖励学习(即RLHF)是一种允许非专家用户通过对智能体行为的短片段表达偏好来传达复杂目标的方法。尽管这种参数化方法能够编码基础任务中存在的复杂隐性知识,但它隐含假设人类无法提供比二元偏好标签更丰富的反馈,导致反馈复杂度过高且用户体验不佳。虽然提供目标详细符号闭式规范可能具有吸引力,但即使对专家用户而言这并非总是可行。然而,在大多数情况下,人类能够意识到智能体应如何沿有意义的行为轴调整其行为以实现潜在目标,即使他们无法完全符号化地指定任务目标。基于这一动机,我们引入了相对行为属性的概念,允许用户通过符号概念(例如增加智能体移动的柔软度或速度)调整智能体行为。我们提出了两种实用方法,能够从有序行为片段中学习建模任何类型的行为属性。我们在包含九种不同行为属性的四项任务上验证了方法的有效性,表明一旦属性被学习,最终用户只需提供约十次反馈即可相对轻松地生成期望的智能体行为。这比流行的基于人类偏好学习的基线方法所需的反馈量低一个数量级以上。补充视频和源代码可在以下链接获取:https://guansuns.github.io/pages/rba。