Understanding and manipulating deformable objects (e.g., ropes and fabrics) is an essential yet challenging task with broad applications. Difficulties come from complex states and dynamics, diverse configurations and high-dimensional action space of deformable objects. Besides, the manipulation tasks usually require multiple steps to accomplish, and greedy policies may easily lead to local optimal states. Existing studies usually tackle this problem using reinforcement learning or imitating expert demonstrations, with limitations in modeling complex states or requiring hand-crafted expert policies. In this paper, we study deformable object manipulation using dense visual affordance, with generalization towards diverse states, and propose a novel kind of foresightful dense affordance, which avoids local optima by estimating states' values for long-term manipulation. We propose a framework for learning this representation, with novel designs such as multi-stage stable learning and efficient self-supervised data collection without experts. Experiments demonstrate the superiority of our proposed foresightful dense affordance. Project page: https://hyperplane-lab.github.io/DeformableAffordance
翻译:理解和操作可变形物体(如绳索和织物)是一项重要但具有挑战性的任务,具有广泛的应用前景。其难点在于可变形物体复杂的状态与动力学、多样化的构型以及高维动作空间。此外,操作任务通常需要多步完成,贪婪策略易导致局部最优状态。现有研究通常采用强化学习或模仿专家示范来解决该问题,但存在对复杂状态建模有限或需要手工设计专家策略的局限。本文研究了基于密集视觉可操作性实现可变形物体操作的方法,并具备对多样化状态的泛化能力;我们提出了一种新型的预见性密集可操作性,通过评估用于长期操作的状态值来避免局部最优。我们设计了一个学习该表征的框架,包含多阶段稳定学习和无需专家的高效自监督数据收集等创新设计。实验证明了所提预见性密集可操作性的优越性。项目页面:https://hyperplane-lab.github.io/DeformableAffordance