Grasping small objects surrounded by unstable or non-rigid material plays a crucial role in applications such as surgery, harvesting, construction, disaster recovery, and assisted feeding. This task is especially difficult when fine manipulation is required in the presence of sensor noise and perception errors; errors inevitably trigger dynamic motion, which is challenging to model precisely. Circumventing the difficulty to build accurate models for contacts and dynamics, data-driven methods like reinforcement learning (RL) can optimize task performance via trial and error, reducing the need for accurate models of contacts and dynamics. Applying RL methods to real robots, however, has been hindered by factors such as prohibitively high sample complexity or the high training infrastructure cost for providing resets on hardware. This work presents CherryBot, an RL system that uses chopsticks for fine manipulation that surpasses human reactiveness for some dynamic grasping tasks. By integrating imprecise simulators, suboptimal demonstrations and external state estimation, we study how to make a real-world robot learning system sample efficient and general while reducing the human effort required for supervision. Our system shows continual improvement through 30 minutes of real-world interaction: through reactive retry, it achieves an almost 100% success rate on the demanding task of using chopsticks to grasp small objects swinging in the air. We demonstrate the reactiveness, robustness and generalizability of CherryBot to varying object shapes and dynamics (e.g., external disturbances like wind and human perturbations). Videos are available at https://goodcherrybot.github.io/.
翻译:在外部不稳定或非刚性材料包围的环境中抓取小物体,在手术、采收、建筑施工、灾后救援及辅助进食等应用中至关重要。当需要精细操作且存在传感器噪声和感知误差时,这项任务尤为困难——误差必然引发动态运动,而这类运动难以精确建模。为规避建立接触与动力学精确模型的难点,强化学习等数据驱动方法可通过试错优化任务性能,减少对接触和动力学精确模型的需求。然而,将强化学习方法应用于真实机器人一直受限于样本复杂度过高或提供硬件复位的高昂训练基础设施成本等障碍。本文提出的CherryBot系统使用筷子进行精细操作,在某些动态抓取任务中超越了人类的反应能力。通过整合非精确仿真器、次优示范和外部状态估计,我们研究了如何使真实世界机器人学习系统兼具样本高效性和泛化性,同时降低所需的人工监督成本。该系统通过30分钟的真实世界交互实现持续改进:借助反应式重试机制,在使用筷子抓取空中摆动小物体这一高难度任务中取得了近乎100%的成功率。我们验证了CherryBot对多样化物体形状和动态特性(如风和人为扰动等外部干扰)的反应性、鲁棒性和泛化能力。视频见https://goodcherrybot.github.io/。