This paper presents an efficient approach to object manipulation planning using Monte Carlo Tree Search (MCTS) to find contact sequences and an efficient ADMM-based trajectory optimization algorithm to evaluate the dynamic feasibility of candidate contact sequences. To accelerate MCTS, we propose a methodology to learn a goal-conditioned policy-value network to direct the search towards promising nodes. Further, manipulation-specific heuristics enable to drastically reduce the search space. Systematic object manipulation experiments in a physics simulator and on real hardware demonstrate the efficiency of our approach. In particular, our approach scales favorably for long manipulation sequences thanks to the learned policy-value network, significantly improving planning success rate.
翻译:本文提出了一种基于蒙特卡洛树搜索(MCTS)的高效物体操作规划方法,用于生成接触序列,并采用基于ADMM的高效轨迹优化算法评估候选接触序列的动力学可行性。为加速MCTS,我们提出了一种学习目标导向策略-价值网络的方案,引导搜索朝向高潜力节点。此外,面向操作任务的专用启发式策略显著缩减了搜索空间。在物理仿真器与真实硬件上的系统性物体操作实验验证了本方法的高效性。特别地,得益于所学习的策略-价值网络,本方法在长序列操作任务中表现出良好的可扩展性,显著提升了规划成功率。