For robots to be effectively deployed in novel environments and tasks, they must be able to understand the feedback expressed by humans during intervention. This can either correct undesirable behavior or indicate additional preferences. Existing methods either require repeated episodes of interactions or assume prior known reward features, which is data-inefficient and can hardly transfer to new tasks. We relax these assumptions by describing human tasks in terms of object-centric sub-tasks and interpreting physical interventions in relation to specific objects. Our method, Object Preference Adaptation (OPA), is composed of two key stages: 1) pre-training a base policy to produce a wide variety of behaviors, and 2) online-updating according to human feedback. The key to our fast, yet simple adaptation is that general interaction dynamics between agents and objects are fixed, and only object-specific preferences are updated. Our adaptation occurs online, requires only one human intervention (one-shot), and produces new behaviors never seen during training. Trained on cheap synthetic data instead of expensive human demonstrations, our policy correctly adapts to human perturbations on realistic tasks on a physical 7DOF robot. Videos, code, and supplementary material are provided.
翻译:为了使机器人能够有效部署于新环境和新任务中,它们必须理解人类在干预过程中表达的反馈。这些反馈既能纠正不受欢迎的行为,也能指示额外的偏好。现有方法要么需要多次交互回合,要么假设预先已知奖励特征,导致数据效率低下且难以迁移至新任务。我们通过以物体为中心的子任务描述人类任务,并解释与特定物体相关的物理干预,从而放宽了这些假设。我们的方法——物体偏好适应(OPA)——由两个关键阶段组成:1) 预训练一个基础策略以产生多种多样的行为,2) 根据人类反馈进行在线更新。我们快速且简单的适应过程的关键在于,智能体与物体之间的一般交互动力学是固定的,仅更新与物体相关的特定偏好。我们的适应过程在线进行,仅需一次人类干预(单次),并能产生训练中从未见过的新行为。模型使用廉价的合成数据而非昂贵的人类演示进行训练,能够在物理七自由度机器人上针对现实任务正确适应人类的扰动。我们提供了视频、代码和补充材料。