The empirical success of Reinforcement Learning (RL) in the setting of contact-rich manipulation leaves much to be understood from a model-based perspective, where the key difficulties are often attributed to (i) the explosion of contact modes, (ii) stiff, non-smooth contact dynamics and the resulting exploding / discontinuous gradients, and (iii) the non-convexity of the planning problem. The stochastic nature of RL addresses (i) and (ii) by effectively sampling and averaging the contact modes. On the other hand, model-based methods have tackled the same challenges by smoothing contact dynamics analytically. Our first contribution is to establish the theoretical equivalence of the two methods for simple systems, and provide qualitative and empirical equivalence on a number of complex examples. In order to further alleviate (ii), our second contribution is a convex, differentiable and quasi-dynamic formulation of contact dynamics, which is amenable to both smoothing schemes, and has proven through experiments to be highly effective for contact-rich planning. Our final contribution resolves (iii), where we show that classical sampling-based motion planning algorithms can be effective in global planning when contact modes are abstracted via smoothing. Applying our method on a collection of challenging contact-rich manipulation tasks, we demonstrate that efficient model-based motion planning can achieve results comparable to RL with dramatically less computation. Video: https://youtu.be/12Ew4xC-VwA
翻译:强化学习(RL)在富接触操作场景中的经验成功,从基于模型的角度仍有诸多未解之谜,其关键困难常被归因于:(i) 接触模式的爆炸式增长,(ii) 刚性、非光滑的接触动力学及其引发的梯度爆炸/不连续问题,以及(iii) 规划问题的非凸性。RL的随机性通过有效采样和平均接触模式来处理(i)和(ii)。另一方面,基于模型的方法通过分析性平滑接触动力学应对同样挑战。我们的第一个贡献是为简单系统建立两种方法的理论等价性,并在多个复杂示例上证明其定性与经验等效性。为进一步缓解(ii),第二个贡献是提出一种凸性、可微且准动态的接触动力学公式,该公式兼容两类平滑方案,并通过实验证明对富接触规划高度有效。最终贡献解决了(iii):我们证明当接触模式通过平滑抽象后,经典基于采样的运动规划算法在全局规划中同样有效。将所提方法应用于一系列挑战性富接触操作任务,我们展示了高效基于模型的运动规划能以显著更少的计算量取得与RL相当的结果。视频链接:https://youtu.be/12Ew4xC-VwA