Aligning agent behaviors with diverse human preferences remains a challenging problem in reinforcement learning (RL), owing to the inherent abstractness and mutability of human preferences. To address these issues, we propose AlignDiff, a novel framework that leverages RL from Human Feedback (RLHF) to quantify human preferences, covering abstractness, and utilizes them to guide diffusion planning for zero-shot behavior customizing, covering mutability. AlignDiff can accurately match user-customized behaviors and efficiently switch from one to another. To build the framework, we first establish the multi-perspective human feedback datasets, which contain comparisons for the attributes of diverse behaviors, and then train an attribute strength model to predict quantified relative strengths. After relabeling behavioral datasets with relative strengths, we proceed to train an attribute-conditioned diffusion model, which serves as a planner with the attribute strength model as a director for preference aligning at the inference phase. We evaluate AlignDiff on various locomotion tasks and demonstrate its superior performance on preference matching, switching, and covering compared to other baselines. Its capability of completing unseen downstream tasks under human instructions also showcases the promising potential for human-AI collaboration. More visualization videos are released on https://aligndiff.github.io/.
翻译:在强化学习中,将智能体行为与多样化人类偏好对齐仍是一个挑战性问题,这源于人类偏好固有的抽象性和易变性。为解决这些问题,我们提出AlignDiff这一全新框架,该框架利用人类反馈强化学习(RLHF)对抽象性偏好进行量化,并指导扩散规划实现零样本行为定制以应对易变性。AlignDiff能精准匹配用户定制化行为,并高效实现行为切换。为构建该框架,我们首先建立多视角人类反馈数据集,其中包含多样化行为属性的比较信息,随后训练属性强度模型以预测量化相对强度。在利用相对强度重新标注行为数据集后,我们训练一个属性条件扩散模型,该模型作为规划器,并结合作为导向器的属性强度模型在推理阶段进行偏好对齐。我们在多种运动任务上评估AlignDiff,结果表明其在偏好匹配、切换和覆盖方面均优于其他基线方法。该方法在人类指令下完成未见下游任务的能力,也展示了人机协作的巨大潜力。更多可视化视频已发布于https://aligndiff.github.io/。