Extracting structured representations from raw visual data is an important and long-standing challenge in machine learning. Recently, techniques for unsupervised learning of object-centric representations have raised growing interest. In this context, enhancing the robustness of the latent features can improve the efficiency and effectiveness of the training of downstream tasks. A promising step in this direction is to disentangle the factors that cause variation in the data. Previously, Invariant Slot Attention disentangled position, scale, and orientation from the remaining features. Extending this approach, we focus on separating the shape and texture components. In particular, we propose a novel architecture that biases object-centric models toward disentangling shape and texture components into two non-overlapping subsets of the latent space dimensions. These subsets are known a priori, hence before the training process. Experiments on a range of object-centric benchmarks reveal that our approach achieves the desired disentanglement while also numerically improving baseline performance in most cases. In addition, we show that our method can generate novel textures for a specific object or transfer textures between objects with distinct shapes.
翻译:从原始视觉数据中提取结构化表示是机器学习领域一个长期存在的重要挑战。近年来,无监督学习目标中心表示的技术引发了越来越多的关注。在此背景下,增强潜在特征的鲁棒性能够提升下游任务训练的效率和效果。该方向上一个有前景的步骤是解耦导致数据变化的因素。此前,不变槽注意力(Invariant Slot Attention)将位置、尺度和方向与其余特征解耦。我们扩展了这一方法,专注于分离形状和纹理成分。具体而言,我们提出了一种新颖的架构,该架构促使目标中心模型将形状和纹理成分解耦到潜在空间维度中两个互不重叠的子集内。这些子集是先验已知的,即在训练过程之前就已确定。在一系列目标中心基准实验上的结果表明,我们的方法在大多情况下实现了预期的解耦效果,同时在数值上改进了基线性能。此外,我们证明了该方法能够为特定对象生成新颖纹理,或在形状不同的对象之间迁移纹理。