Generalizing from individual skill executions to long-horizon tasks is a core challenge in building autonomous robots. A promising direction is learning high-level, symbolic representations of low-level robot skills, enabling abstract reasoning independent of the low-level state space. Recent advances in foundation models have made it possible to generate symbolic predicates that operate on raw sensory inputs-a process we call generative predicate invention-to facilitate downstream representation learning. However, prior work learns these abstractions using heuristic or ad-hoc procedures, ignoring the question of which formal properties they ought to satisfy, and how to guarantee these properties. We address these questions by presenting a formal theory of generative predicate invention for task-level planning, and proposing SkillWrapper, a method that learns symbolic models for provably sound and complete planning. Our approach leverages foundation models to actively collect robot data and learn human-interpretable, plannable representations, using only RGB image observations. Our extensive empirical evaluation in simulation and on real robots shows that SkillWrapper learns abstract representations that enable robots to compose black-box skills to solve unseen, long-horizon tasks in the real world.
翻译:摘要:从个体技能执行泛化到长期任务,是构建自主机器人的核心挑战之一。一个有前景的方向是学习高级、符号化的底层机器人技能表征,从而实现对低层状态空间无关的抽象推理。基础模型的最新进展使得生成基于原始感官输入的符号谓词成为可能——我们称此过程为生成式谓词发明——以促进下游表征学习。然而,先前的工作采用启发式或特设程序学习这些抽象表征,忽略了它们应满足哪些形式化属性以及如何保证这些属性的问题。我们通过提出用于任务级规划的生成式谓词发明形式化理论,并设计SkillWrapper方法来解决这些问题。该方法学习可证明完备且可靠的规划符号模型,利用基础模型主动收集机器人数据,学习仅基于RGB图像观测的人类可解释、可规划表征。我们在仿真和真实机器人上开展的广泛实证评估表明,SkillWrapper所学习的抽象表征能使机器人组合黑盒技能,解决真实世界中未见过的长期任务。