Objects rarely sit in isolation in everyday human environments. If we want robots to operate and perform tasks in our human environments, they must understand how the objects they manipulate will interact with structural elements of the environment for all but the simplest of tasks. As such, we'd like our robots to reason about how multiple objects and environmental elements relate to one another and how those relations may change as the robot interacts with the world. We examine the problem of predicting inter-object and object-environment relations between previously unseen objects and novel environments purely from partial-view point clouds. Our approach enables robots to plan and execute sequences to complete multi-object manipulation tasks defined from logical relations. This removes the burden of providing explicit, continuous object states as goals to the robot. We explore several different neural network architectures for this task. We find the best performing model to be a novel transformer-based neural network that both predicts object-environment relations and learns a latent-space dynamics function. We achieve reliable sim-to-real transfer without any fine-tuning. Our experiments show that our model understands how changes in observed environmental geometry relate to semantic relations between objects. We show more videos on our website: https://sites.google.com/view/erelationaldynamics.
翻译:在日常人类环境中,物体很少孤立存在。若要让机器人在我们的环境中执行操作任务,除最简单的任务外,它们必须理解所操控物体与环境结构元素之间的相互作用方式。因此,我们希望机器人能够推理多物体与环境要素之间的关联性,以及这些关系在机器人与环境交互时可能发生的变化。我们研究仅从局部视角点云预测未见过物体与新颖环境之间的物体间关系和物体-环境关系的问题。所提方法使机器人能够基于逻辑关系定义的多物体操作任务,规划并执行完整序列,从而免除了向机器人提供显式连续物体状态作为目标的负担。我们探索了多种神经网络架构来完成该任务,发现最佳模型是一种新颖的基于Transformer的神经网络——该网络既能预测物体-环境关系,又能学习潜空间动力学函数。我们在无需任何微调的情况下实现了可靠的仿真到现实迁移。实验表明,该模型能够理解观测环境几何结构的变化如何与物体间的语义关系相关联。更多演示视频请访问我们的网站:https://sites.google.com/view/erelationaldynamics。