The binding problem in artificial neural networks is actively explored with the goal of achieving human-level recognition skills through the comprehension of the world in terms of symbol-like entities. Especially in the field of computer vision, object-centric learning (OCL) is extensively researched to better understand complex scenes by acquiring object representations or slots. While recent studies in OCL have made strides with complex images or videos, the interpretability and interactivity over object representation remain largely uncharted, still holding promise in the field of OCL. In this paper, we introduce a novel method, Slot Attention with Image Augmentation (SlotAug), to explore the possibility of learning interpretable controllability over slots in a self-supervised manner by utilizing an image augmentation strategy. We also devise the concept of sustainability in controllable slots by introducing iterative and reversible controls over slots with two proposed submethods: Auxiliary Identity Manipulation and Slot Consistency Loss. Extensive empirical studies and theoretical validation confirm the effectiveness of our approach, offering a novel capability for interpretable and sustainable control of object representations. Code will be available soon.
翻译:人工神经网络中的绑定问题正被积极探索,旨在通过理解以符号化实体呈现的世界,实现人类级别的识别能力。尤其在计算机视觉领域,对象中心学习(OCL)通过获取对象表示或槽位来更好地理解复杂场景,已得到广泛研究。尽管近期OCL研究在复杂图像或视频处理方面取得了进展,但对象表示的可解释性和交互性仍 largely 未被探索,这仍是OCL领域颇具前景的研究方向。本文提出一种新方法——图像增强槽位注意力(Slot Attention with Image Augmentation, SlotAug),通过利用图像增强策略,探索以自监督方式学习槽位可解释可控性的可能性。我们进一步引入可持续可控槽位的概念,通过两种子方法——辅助恒等操作与槽位一致性损失——实现槽位的迭代与可逆控制。大量实验研究与理论验证证明了该方法的有效性,为对象表示的可解释、可持续控制提供了全新能力。代码将近期开源。