Robots operating in human environments must be able to rearrange objects into semantically-meaningful configurations, even if these objects are previously unseen. In this work, we focus on the problem of building physically-valid structures without step-by-step instructions. We propose StructDiffusion, which combines a diffusion model and an object-centric transformer to construct structures given partial-view point clouds and high-level language goals, such as "set the table". Our method can perform multiple challenging language-conditioned multi-step 3D planning tasks using one model. StructDiffusion even improves the success rate of assembling physically-valid structures out of unseen objects by on average 16% over an existing multi-modal transformer model trained on specific structures. We show experiments on held-out objects in both simulation and on real-world rearrangement tasks. Importantly, we show how integrating both a diffusion model and a collision-discriminator model allows for improved generalization over other methods when rearranging previously-unseen objects. For videos and additional results, see our website: https://structdiffusion.github.io/.
翻译:机器人在人类环境中操作时,必须能够将物体重新排列成具有语义意义的配置,即使这些物体此前未被见过。本文聚焦于在没有逐步指令的情况下构建物理有效结构的问题。我们提出StructDiffusion,该方法结合扩散模型和以物体为中心的Transformer,根据部分视角点云和高级语言目标(如“摆好餐桌”)构建结构。我们的方法能够使用单一模型执行多项具有挑战性的、基于语言条件的多步3D规划任务。StructDiffusion甚至能够在利用未见物体组装物理有效结构时,平均成功率比现有针对特定结构训练的多模态Transformer模型提高16%。我们在仿真和真实世界重排任务中对保留物体进行了实验。重要的是,我们展示了如何通过整合扩散模型和碰撞判别模型,在重排之前未见过的物体时,相比其他方法实现更好的泛化能力。视频和其他结果请见我们的网站:https://structdiffusion.github.io/。