Object rearrangement is a challenge for embodied agents because solving these tasks requires generalizing across a combinatorially large set of configurations of entities and their locations. Worse, the representations of these entities are unknown and must be inferred from sensory percepts. We present a hierarchical abstraction approach to uncover these underlying entities and achieve combinatorial generalization from unstructured visual inputs. By constructing a factorized transition graph over clusters of entity representations inferred from pixels, we show how to learn a correspondence between intervening on states of entities in the agent's model and acting on objects in the environment. We use this correspondence to develop a method for control that generalizes to different numbers and configurations of objects, which outperforms current offline deep RL methods when evaluated on simulated rearrangement tasks.
翻译:物体重排对于具身代理而言是一项挑战,因为解决这类任务需要泛化到实体及其位置组合构成的海量构型。更困难的是,这些实体的表征是未知的,必须从感知数据中推断得出。我们提出了一种层级抽象方法,用于从非结构化视觉输入中揭示这些潜在实体,并实现组合泛化。通过构建基于像素推断的实体表征簇上的因式化转移图,我们展示了如何学习代理模型中对实体状态的干预与环境中对物体的操作之间的对应关系。利用这种对应关系,我们开发了一种能够泛化到不同物体数量与构型的控制方法——该方法在模拟重排任务评估中显著优于当前的离线深度强化学习方法。