Teaching dexterity to multi-fingered robots has been a longstanding challenge in robotics. Most prominent work in this area focuses on learning controllers or policies that either operate on visual observations or state estimates derived from vision. However, such methods perform poorly on fine-grained manipulation tasks that require reasoning about contact forces or about objects occluded by the hand itself. In this work, we present T-Dex, a new approach for tactile-based dexterity, that operates in two phases. In the first phase, we collect 2.5 hours of play data, which is used to train self-supervised tactile encoders. This is necessary to bring high-dimensional tactile readings to a lower-dimensional embedding. In the second phase, given a handful of demonstrations for a dexterous task, we learn non-parametric policies that combine the tactile observations with visual ones. Across five challenging dexterous tasks, we show that our tactile-based dexterity models outperform purely vision and torque-based models by an average of 1.7X. Finally, we provide a detailed analysis on factors critical to T-Dex including the importance of play data, architectures, and representation learning.
翻译:教导多指机器人掌握灵巧操作一直是机器人学领域的长期挑战。该领域最具代表性的工作主要聚焦于学习基于视觉观测或视觉状态估计的控制器或策略。然而,这类方法在需要推理接触力或被手部遮挡物体的精细操作任务中表现欠佳。本文提出T-Dex——一种基于触觉的灵巧操作新方法,该方法分两个阶段运行。第一阶段,我们收集2.5小时的玩耍数据,用于训练自监督触觉编码器。这一步骤对于将高维触觉读数降至低维嵌入空间至关重要。第二阶段,针对灵巧任务的少量示范,我们学习结合触觉与视觉观测的非参数化策略。在五个具有挑战性的灵巧任务中,我们的基于触觉的灵巧模型平均性能比纯视觉和基于扭矩的模型高出1.7倍。最后,我们详细分析了影响T-Dex的关键因素,包括玩耍数据的重要性、架构设计及表征学习。