Object-centric learning aims to break down complex visual scenes into more manageable object representations, enhancing the understanding and reasoning abilities of machine learning systems toward the physical world. Recently, slot-based video models have demonstrated remarkable proficiency in segmenting and tracking objects, but they overlook the importance of the effective reasoning module. In the real world, reasoning and predictive abilities play a crucial role in human perception and object tracking; in particular, these abilities are closely related to human intuitive physics. Inspired by this, we designed a novel reasoning module called the Slot-based Time-Space Transformer with Memory buffer (STATM) to enhance the model's perception ability in complex scenes. The memory buffer primarily serves as storage for slot information from upstream modules, the Slot-based Time-Space Transformer makes predictions through slot-based spatiotemporal attention computations and fusion. Our experiment results on various datasets show that STATM can significantly enhance object-centric learning capabilities of slot-based video models.
翻译:对象中心学习旨在将复杂视觉场景分解为更易于处理的对象表示,从而提升机器学习系统对物理世界的理解和推理能力。近年来,基于槽的视频模型在对象分割与跟踪方面展现出卓越性能,但它们忽略了有效推理模块的重要性。在现实世界中,推理与预测能力对人类感知和对象跟踪起着关键作用,尤其与人类的直觉物理学密切相关。受此启发,我们设计了一种名为基于槽的时空变换器与记忆缓冲区(STATM)的新型推理模块,以增强模型在复杂场景中的感知能力。记忆缓冲区主要作为上游模块槽信息的存储单元,而基于槽的时空变换器则通过基于槽的时空注意力计算与融合进行预测。我们在多种数据集上的实验结果表明,STATM能够显著提升基于槽的视频模型的对象中心学习能力。