Vague objectives in many real-life scenarios pose long-standing challenges for robotics, as defining rules, rewards, or constraints for optimization is difficult. Tasks like tidying a messy table may appear simple for humans, but articulating the criteria for tidiness is complex due to the ambiguity and flexibility in commonsense reasoning. Recent advancement in Large Language Models (LLMs) offers us an opportunity to reason over these vague objectives: learned from extensive human data, LLMs capture meaningful common sense about human behavior. However, as LLMs are trained solely on language input, they may struggle with robotic tasks due to their limited capacity to account for perception and low-level controls. In this work, we propose a simple approach to solve the task of table tidying, an example of robotic tasks with vague objectives. Specifically, the task of tidying a table involves not just clustering objects by type and functionality for semantic tidiness but also considering spatial-visual relations of objects for a visually pleasing arrangement, termed as visual tidiness. We propose to learn a lightweight, image-based tidiness score function to ground the semantically tidy policy of LLMs to achieve visual tidiness. We innovatively train the tidiness score using synthetic data gathered using random walks from a few tidy configurations. Such trajectories naturally encode the order of tidiness, thereby eliminating the need for laborious and expensive human demonstrations. Our empirical results show that our pipeline can be applied to unseen objects and complex 3D arrangements.
翻译:现实场景中的模糊目标对机器人技术构成长期挑战,因为为其优化定义规则、奖励或约束十分困难。整理凌乱餐桌等任务对人类而言看似简单,但由于常识推理中存在歧义性和灵活性,明确表述"整洁"的标准却十分复杂。大语言模型的最新进展为我们提供了应对模糊目标推理的契机:通过学习海量人类数据,LLMs能够捕获人类行为中有意义的常识。然而,由于LLMs仅基于语言输入进行训练,其在感知和底层控制方面的能力有限,可能难以胜任机器人任务。本文提出一种简洁方法来解决餐桌整理这一典型模糊目标机器人任务。具体而言,餐桌整理不仅需要按类型和功能对物体进行聚类以实现语义整洁,还需考虑物体的空间视觉关系以形成视觉和谐的布局,即视觉整洁。我们提出学习轻量级的基于图像的整洁度评分函数,将LLMs的语义整洁策略落地以实现视觉整洁。创新性地采用从少量整洁配置出发的随机游走生成的合成数据训练该评分函数。这些轨迹自然编码了整洁程度的序关系,从而避免耗费人力且昂贵的人体演示。实验结果表明,我们的流程可应用于未见过的物体及复杂三维布局。