In this paper, we propose a novel affordance model, which combines object, action, and effect information in the latent space of a predictive neural network architecture that is built on Conditional Neural Processes. Our model allows us to make predictions of intermediate effects expected to be obtained during action executions and make multi-step plans that include partial actions. We first compared the prediction capability of our model using an existing interaction data set and showed that it outperforms a recurrent neural network-based model in predicting the effects of lever-up actions. Next, we showed that our model can generate accurate effect predictions for other actions, such as push and grasp actions. Our system was shown to generate successful multi-step plans to bring objects to desired positions using the traditional A* search algorithm. Furthermore, we realized a continuous planning method and showed that the proposed system generated more accurate and effective plans with sequences of partial action executions compared to plans that only consider full action executions using both planning algorithms.
翻译:在本文中,我们提出了一种新颖的affordance模型,该模型在基于条件神经过程的预测性神经网络架构的潜在空间中,融合了物体、动作和效应信息。我们的模型能够预测动作执行过程中预期获得的中间效应,并制定包含部分动作的多步规划。首先,我们利用现有交互数据集对模型的预测能力进行了对比,结果表明,在预测杠杆拉起动作的效应方面,该模型优于基于循环神经网络的模型。其次,我们证明该模型能对其他动作(如推和抓取动作)生成准确的效应预测。实验表明,利用传统A*搜索算法,我们的系统能够生成成功实现物体移动到目标位置的多步规划。此外,我们实现了连续规划方法,并证明与仅考虑完整动作执行的规划相比,采用两种规划算法时,基于部分动作执行序列的所提系统能生成更准确、更有效的规划。