Multiview Self-Supervised Learning (MSSL) is based on learning invariances with respect to a set of input transformations. However, invariance partially or totally removes transformation-related information from the representations, which might harm performance for specific downstream tasks that require such information. We propose 2D strUctured and EquivarianT representations (coined DUET), which are 2d representations organized in a matrix structure, and equivariant with respect to transformations acting on the input data. DUET representations maintain information about an input transformation, while remaining semantically expressive. Compared to SimCLR (Chen et al., 2020) (unstructured and invariant) and ESSL (Dangovski et al., 2022) (unstructured and equivariant), the structured and equivariant nature of DUET representations enables controlled generation with lower reconstruction error, while controllability is not possible with SimCLR or ESSL. DUET also achieves higher accuracy for several discriminative tasks, and improves transfer learning.
翻译:多视角自监督学习(MSSL)基于学习对一组输入变换的不变性。然而,不变性会部分或完全去除表示中的变换相关信息,这可能损害需要此类信息的特定下游任务的性能。我们提出二维结构化且等变表示(命名为DUET),其组织为矩阵结构的二维表示,并对作用于输入数据的变换具有等变性。DUET表示在保持输入变换信息的同时,仍具有语义表达能力。与SimCLR(Chen等,2020年)(非结构化且不变)和ESSL(Dangovski等,2022年)(非结构化且等变)相比,DUET表示的结构化和等变性使得可控生成具有更低的重构误差,而SimCLR或ESSL无法实现可控性。DUET在多个判别任务中取得了更高准确率,并提升了迁移学习性能。