We introduce a novel approach to the executable semantic object rearrangement problem. In this challenge, a robot seeks to create an actionable plan that rearranges objects within a scene according to a pattern dictated by a natural language description. Unlike existing methods such as StructFormer and StructDiffusion, which tackle the issue in two steps by first generating poses and then leveraging a task planner for action plan formulation, our method concurrently addresses pose generation and action planning. We achieve this integration using a Language-Guided Monte-Carlo Tree Search (LGMCTS). Quantitative evaluations are provided on two simulation datasets, and complemented by qualitative tests with a real robot.
翻译:我们提出了一种解决可执行语义对象重排问题的新方法。在该挑战中,机器人需根据自然语言描述所指定的模式,制定可操作的计划以重新排列场景中的物体。与现有方法(如StructFormer和StructDiffusion)分两步处理(首先生成位姿,然后利用任务规划器制定动作计划)不同,我们的方法同时处理位姿生成与动作规划。通过使用语言引导的蒙特卡洛树搜索(LGMCTS)实现了这一集成。我们在两个仿真数据集上提供了定量评估,并通过真实机器人实验补充了定性测试。