Dextrous in-hand manipulation with a multi-fingered robotic hand is a challenging task, esp. when performed with the hand oriented upside down, demanding permanent force-closure, and when no external sensors are used. For the task of reorienting an object to a given goal orientation (vs. infinitely spinning it around an axis), the lack of external sensors is an additional fundamental challenge as the state of the object has to be estimated all the time, e.g., to detect when the goal is reached. In this paper, we show that the task of reorienting a cube to any of the 24 possible goal orientations in a ${\pi}$/2-raster using the torque-controlled DLR-Hand II is possible. The task is learned in simulation using a modular deep reinforcement learning architecture: the actual policy has only a small observation time window of 0.5s but gets the cube state as an explicit input which is estimated via a deep differentiable particle filter trained on data generated by running the policy. In simulation, we reach a success rate of 92% while applying significant domain randomization. Via zero-shot Sim2Real-transfer on the real robotic system, all 24 goal orientations can be reached with a high success rate.
翻译:使用多指机械手进行灵巧手内操作是一项极具挑战性的任务,尤其在手掌朝下(即倒置姿态)操作时,需要维持恒定的力闭合,且不借助外部传感器。对于将物体重定向至指定目标姿态(而非使其绕轴无限旋转)的任务而言,缺乏外部传感器还带来一个根本性难题:必须持续估计物体状态(例如检测目标是否达成)。本文证明,利用力矩控制的DLR-Hand II机械手,将立方体重定向至${\pi}/2$栅格内24种可能目标姿态中的任意一种具有可行性。该任务通过模块化深度强化学习架构在仿真环境中习得:实际策略仅需0.5秒的短时观测窗口,但通过显式输入经由深度可微粒子滤波器(该滤波器基于策略运行生成的数据训练)估计的立方体状态。在仿真中,我们通过大量领域随机化达到了92%的成功率。通过在真实机器人系统上零样本Sim2Real迁移,所有24种目标姿态均能以高成功率达成。