Generating human-like behavior on robots is a great challenge especially in dexterous manipulation tasks with robotic hands. Even in simulation with no sample constraints, scripting controllers is intractable due to high degrees of freedom, and manual reward engineering can also be hard and lead to non-realistic motions. Leveraging the recent progress on Reinforcement Learning from Human Feedback (RLHF), we propose a framework to learn a universal human prior using direct human preference feedback over videos, for efficiently tuning the RL policy on 20 dual-hand robot manipulation tasks in simulation, without a single human demonstration. One task-agnostic reward model is trained through iteratively generating diverse polices and collecting human preference over the trajectories; it is then applied for regularizing the behavior of polices in the fine-tuning stage. Our method empirically demonstrates more human-like behaviors on robot hands in diverse tasks including even unseen tasks, indicating its generalization capability.
翻译:在机器人上生成类人行为是一个巨大挑战,尤其在机械手灵巧操作任务中。即使在无样本限制的仿真环境中,由于自由度较高,脚本控制器难以实现,手动奖励工程同样困难且易产生非真实运动。借助从人类反馈中强化学习(RLHF)的最新进展,我们提出一个框架,通过直接收集视频上的人类偏好反馈来学习通用人类先验,从而高效调节20种双手机器人灵巧操作任务的强化学习策略(在仿真环境中,且无需任何人类示范)。通过迭代生成多样化策略并收集轨迹上的人类偏好,训练出一个任务无关的奖励模型;该模型随后用于微调阶段对策略行为进行正则化。实验表明,我们的方法在包括未见任务在内的多样化任务中,能使机械手展现出更类人行为,验证了其泛化能力。