Revolutionizing the field of deep learning, Transformer-based models have achieved remarkable performance in many tasks. Recent research has recognized these models are robust to shuffling but are limited to inter-token permutation in the forward propagation. In this work, we propose our definition of permutation equivariance, a broader concept covering both inter- and intra- token permutation in the forward and backward propagation of neural networks. We rigorously proved that such permutation equivariance property can be satisfied on most vanilla Transformer-based models with almost no adaptation. We examine the property over a range of state-of-the-art models including ViT, Bert, GPT, and others, with experimental validations. Further, as a proof-of-concept, we explore how real-world applications including privacy-enhancing split learning, and model authorization, could exploit the permutation equivariance property, which implicates wider, intriguing application scenarios.
翻译:革新深度学习领域的Transformer模型在许多任务中取得了显著性能。近期研究表明这类模型具有抗乱序特性,但仅限于前向传播中的令牌间排列。本研究提出更广义的排列等变性定义,涵盖神经网络前向与反向传播中令牌间及令牌内部的排列操作。我们严格证明了该排列等变性在多数原生Transformer模型上几乎无需适配即可实现。通过涵盖ViT、Bert、GPT等前沿模型及其实验验证,我们考查了这一特性。此外,作为概念验证,我们探索了隐私增强型分割学习及模型授权等实际应用如何利用排列等变性,揭示了更广泛的潜在应用场景。