Manipulation of elastoplastic objects like dough often involves topological changes such as splitting and merging. The ability to accurately predict these topological changes that a specific action might incur is critical for planning interactions with elastoplastic objects. We present DoughNet, a Transformer-based architecture for handling these challenges, consisting of two components. First, a denoising autoencoder represents deformable objects of varying topology as sets of latent codes. Second, a visual predictive model performs autoregressive set prediction to determine long-horizon geometrical deformation and topological changes purely in latent space. Given a partial initial state and desired manipulation trajectories, it infers all resulting object geometries and topologies at each step. DoughNet thereby allows to plan robotic manipulation; selecting a suited tool, its pose and opening width to recreate robot- or human-made goals. Our experiments in simulated and real environments show that DoughNet is able to significantly outperform related approaches that consider deformation only as geometrical change.
翻译:处理面团等弹塑性物体的操作常涉及分割与合并等拓扑变化。准确预测特定操作可能引发的拓扑变化,对于规划与弹塑性物体的交互至关重要。我们提出DoughNet——一种基于Transformer架构的解决方案,包含两个组件:首先,去噪自编码器将具有不同拓扑结构的可变形物体表示为潜代码集合;其次,视觉预测模型通过自回归集合预测,在纯潜空间中确定长时域几何形变与拓扑变化。基于部分初始状态与期望操作轨迹,该模型可推断每个步骤中所有结果物体的几何结构与拓扑形态。DoughNet因此能够规划机器人操作:通过选择合适工具、位姿及开口宽度,复现机器人或人类设定的目标。我们的模拟与真实环境实验表明,相比仅将形变视为几何变化的相关方法,DoughNet的性能显著更优。