Interactive imitation learning is an efficient, model-free method through which a robot can learn a task by repetitively iterating an execution of a learning policy and a data collection by querying human demonstrations. However, deploying unmatured policies for clearance-limited tasks, like industrial insertion, poses significant collision risks. For such tasks, a robot should detect the collision risks and request intervention by ceding control to a human when collisions are imminent. The former requires an accurate model of the environment, a need that significantly limits the scope of IIL applications. In contrast, humans implicitly demonstrate environmental precision by adjusting their behavior to avoid collisions when performing tasks. Inspired by human behavior, this paper presents a novel interactive learning method that uses demonstrator-perceived precision as a criterion for human intervention called Demonstrator-perceived Precision-aware Interactive Imitation Learning (DPIIL). DPIIL captures precision by observing the speed-accuracy trade-off exhibited in human demonstrations and cedes control to a human to avoid collisions in states where high precision is estimated. DPIIL improves the safety of interactive policy learning and ensures efficiency without explicitly providing precise information of the environment. We assessed DPIIL's effectiveness through simulations and real-robot experiments that trained a UR5e 6-DOF robotic arm to perform assembly tasks. Our results significantly improved training safety, and our best performance compared favorably with other learning methods.
翻译:交互式模仿学习是一种高效的无模型方法,机器人通过重复执行学习策略并查询人类演示来收集数据,从而习得任务。然而,将未成熟策略部署到高间隙限制任务(如工业插入)中会带来显著的碰撞风险。对于此类任务,机器人需要检测碰撞风险,并在碰撞即将发生时通过将控制权移交给人来请求干预。前者需要精确的环境模型,这极大地限制了交互式模仿学习应用的范围。相比之下,人类在执行任务时会通过调整行为来避免碰撞,从而隐式地展示环境精度。受人类行为启发,本文提出一种新的交互式学习方法,将演示者感知精度作为人类干预的标准,称为演示者感知精度感知的交互式模仿学习。该方法通过观察人类演示中展现的速度-精度权衡来捕获精度,并在估计出高精度的状态下将控制权移交给人类以避免碰撞。DPIIL提升了交互式策略学习的安全性,无需明确提供环境的精确信息即可确保效率。我们通过训练UR5e六自由度机械臂执行装配任务的仿真和真实机器人实验评估了DPIIL的有效性。结果显著改善了训练安全性,我们的最佳性能优于其他学习方法。