We present a novel framework to overcome the limitations of equivariant architectures in learning functions with group symmetries. In contrary to equivariant architectures, we use an arbitrary base model (such as an MLP or a transformer) and symmetrize it to be equivariant to the given group by employing a small equivariant network that parameterizes the probabilistic distribution underlying the symmetrization. The distribution is end-to-end trained with the base model which can maximize performance while reducing sample complexity of symmetrization. We show that this approach ensures not only equivariance to given group but also universal approximation capability in expectation. We implement our method on a simple patch-based transformer that can be initialized from pretrained vision transformers, and test it for a wide range of symmetry groups including permutation and Euclidean groups and their combinations. Empirical tests show competitive results against tailored equivariant architectures, suggesting the potential for learning equivariant functions for diverse groups using a non-equivariant universal base architecture. We further show evidence of enhanced learning in symmetric modalities, like graphs, when pretrained from non-symmetric modalities, like vision. Our implementation will be open-sourced at https://github.com/jw9730/lps.
翻译:我们提出了一种新颖框架,用于克服等变架构在学习具有群对称性的函数时的局限性。与等变架构不同,我们使用任意基础模型(如MLP或Transformer),并通过一个参数化对称化概率分布的小型等变网络对其进行对称化,使其对给定群具有等变性。该分布与基础模型进行端到端训练,可在最大化性能的同时降低对称化的样本复杂度。我们证明,该方法不仅能确保对给定群的等变性,还能在期望意义上实现通用近似能力。我们在一个基于补丁的简单Transformer上实现该方法,该模型可从预训练视觉Transformer初始化,并在包括置换群、欧几里得群及其组合在内的广泛对称群上进行测试。经验测试表明,该方法在竞争性结果上媲美定制等变架构,展现出使用非等变通用基础架构学习不同群等变函数的潜力。此外,我们还展示了从非对称模态(如视觉)预训练后,在对称模态(如图)中增强学习的证据。我们将开源实现代码于https://github.com/jw9730/lps。