We introduce Cosmos, a framework for object-centric world modeling that is designed for compositional generalization (CG), i.e., high performance on unseen input scenes obtained through the composition of known visual "atoms." The central insight behind Cosmos is the use of a novel form of neurosymbolic grounding. Specifically, the framework introduces two new tools: (i) neurosymbolic scene encodings, which represent each entity in a scene using a real vector computed using a neural encoder, as well as a vector of composable symbols describing attributes of the entity, and (ii) a neurosymbolic attention mechanism that binds these entities to learned rules of interaction. Cosmos is end-to-end differentiable; also, unlike traditional neurosymbolic methods that require representations to be manually mapped to symbols, it computes an entity's symbolic attributes using vision-language foundation models. Through an evaluation that considers two different forms of CG on an established blocks-pushing domain, we show that the framework establishes a new state-of-the-art for CG in world modeling.
翻译:我们引入了 Cosmos,这是一个面向对象的世界建模框架,专为组合泛化(CG,即对通过已知视觉“原子”组合而成的未见输入场景实现高性能)而设计。Cosmos 的核心思想在于采用一种新颖的神经符号基础化形式。具体而言,该框架引入了两种新工具:(i)神经符号场景编码,它使用通过神经编码器计算的实向量以及描述实体属性的可组合符号向量来表示场景中的每个实体;(ii)一种神经符号注意力机制,将这些实体与学习到的交互规则绑定起来。Cosmos 是端到端可微的;此外,与需要将表示手动映射为符号的传统神经符号方法不同,它利用视觉-语言基础模型计算实体的符号属性。通过在一个既定的积木推送领域上对两种不同形式的 CG 进行评估,我们证明了该框架在世界建模中的 CG 方面达到了新的最先进水平。