Autonomous excavation is a challenging task. The unknown contact dynamics between the excavator bucket and the terrain could easily result in large contact forces and jamming problems during excavation. Traditional model-based methods struggle to handle such problems due to complex dynamic modeling. In this paper, we formulate the excavation skills with three novel manipulation primitives. We propose to learn the manipulation primitives with offline reinforcement learning (RL) to avoid large amounts of online robot interactions. The proposed method can learn efficient penetration skills from sub-optimal demonstrations, which contain sub-trajectories that can be ``stitched" together to formulate an optimal trajectory without causing jamming. We evaluate the proposed method with extensive experiments on excavating a variety of rigid objects and demonstrate that the learned policy outperforms the demonstrations. We also show that the learned policy can quickly adapt to unseen and challenging fragmented rocks with online fine-tuning.
翻译:自主挖掘是一项具有挑战性的任务。挖掘机铲斗与地形之间未知的接触动力学容易导致挖掘过程中产生较大的接触力和卡堵问题。传统的基于模型的方法由于复杂的动力学建模而难以处理此类问题。本文通过三种新颖的操作基元来构建挖掘技能。我们提出使用离线强化学习来学习这些操作基元,以避免大量在线机器人交互。所提出的方法能够从次优演示中学习高效的穿透技能,这些演示包含可“拼接”的子轨迹,以形成无卡堵的最优轨迹。我们通过大量挖掘各种刚性物体的实验评估了所提出的方法,并证明学习到的策略优于演示。我们还展示了学习到的策略能够通过在线微调快速适应未见过的、具有挑战性的破碎岩石。