The binding problem in artificial neural networks is actively explored with the goal of achieving human-level recognition skills through the comprehension of the world in terms of symbol-like entities. Especially in the field of computer vision, object-centric learning (OCL) is extensively researched to better understand complex scenes by acquiring object representations or slots. While recent studies in OCL have made strides with complex images or videos, the interpretability and interactivity over object representation remain largely uncharted, still holding promise in the field of OCL. In this paper, we introduce a novel method, Slot Attention with Image Augmentation (SlotAug), to explore the possibility of learning interpretable controllability over slots in a self-supervised manner by utilizing an image augmentation strategy. We also devise the concept of sustainability in controllable slots by introducing iterative and reversible controls over slots with two proposed submethods: Auxiliary Identity Manipulation and Slot Consistency Loss. Extensive empirical studies and theoretical validation confirm the effectiveness of our approach, offering a novel capability for interpretable and sustainable control of object representations. Code will be available soon.
翻译:人工神经网络中的绑定问题正被积极探索,旨在通过以符号类实体理解世界,实现人类级别的识别能力。尤其在计算机视觉领域,以对象为中心的学习通过获取对象表征或槽,被广泛研究以更好地理解复杂场景。尽管近期以对象为中心的学习研究在复杂图像或视频方面取得了进展,但对象表征的可解释性和交互性仍基本未受关注,在以对象为中心的学习领域仍有潜力。本文提出了一种新方法——基于图像增强的槽注意力机制,通过利用图像增强策略,探索以自监督方式学习槽的可解释可控性的可能性。我们还通过引入迭代和可逆控制,并借助两种子方法——辅助身份操作与槽一致性损失,提出了可控槽的可持续性概念。大量的实证研究和理论验证证实了我们的方法有效性,为对象表征的可解释且可持续控制提供了新能力。代码即将公开。