An effective approach to solving long-horizon tasks in robotics domains with continuous state and action spaces is bilevel planning, wherein a high-level search over an abstraction of an environment is used to guide low-level decision-making. Recent work has shown how to enable such bilevel planning by learning abstract models in the form of symbolic operators and neural samplers. In this work, we show that existing symbolic operator learning approaches fall short in many robotics domains where a robot's actions tend to cause a large number of irrelevant changes in the abstract state. This is primarily because they attempt to learn operators that exactly predict all observed changes in the abstract state. To overcome this issue, we propose to learn operators that 'choose what to predict' by only modelling changes necessary for abstract planning to achieve specified goals. Experimentally, we show that our approach learns operators that lead to efficient planning across 10 different hybrid robotics domains, including 4 from the challenging BEHAVIOR-100 benchmark, while generalizing to novel initial states, goals, and objects.
翻译:解决具有连续状态和动作空间的机器人领域中的长时域任务的一种有效方法是双层规划,即通过对环境抽象的高层搜索来指导底层决策。最近的研究展示了如何通过学习符号化操作符和神经采样器形式的抽象模型来实现这种双层规划。在本工作中,我们表明,现有的符号化操作符学习方法在许多机器人领域中表现不足,因为机器人的动作往往会导致抽象状态中出现大量无关变化。这主要是因为它们试图学习能精确预测抽象状态中所有观察到的变化的操作符。为克服这一问题,我们提出学习一种“选择预测内容”的操作符,仅对实现指定目标所需的抽象规划变化进行建模。实验表明,我们的方法在10个不同的混合机器人领域(包括来自具有挑战性的BEHAVIOR-100基准的4个领域)中学习了能实现高效规划的操作符,同时能泛化到新的初始状态、目标和物体。