Emerging from the pairwise attention in conventional Transformers, there is a growing interest in sparse attention mechanisms that align more closely with localized, contextual learning in the biological brain. Existing studies such as the Coordination method employ iterative cross-attention mechanisms with a bottleneck to enable the sparse association of inputs. However, these methods are parameter inefficient and fail in more complex relational reasoning tasks. To this end, we propose Associative Transformer (AiT) to enhance the association among sparsely attended input patches, improving parameter efficiency and performance in relational reasoning tasks. AiT leverages a learnable explicit memory, comprised of various specialized priors, with a bottleneck attention to facilitate the extraction of diverse localized features. Moreover, we propose a novel associative memory-enabled patch reconstruction with a Hopfield energy function. The extensive experiments in four image classification tasks with three different sizes of AiT demonstrate that AiT requires significantly fewer parameters and attention layers while outperforming Vision Transformers and a broad range of sparse Transformers. Additionally, AiT establishes new SOTA performance in the Sort-of-CLEVR dataset, outperforming the previous Coordination method.
翻译:从传统Transformer中的成对注意力机制出发,稀疏注意力机制日益受到关注,这类机制更贴近生物大脑中局部化、情境化的学习模式。现有研究(如协调方法)采用带瓶颈的迭代交叉注意力机制来实现输入的稀疏关联。然而,这些方法参数效率低下,在更复杂的关联推理任务中表现不佳。为此,我们提出关联Transformer(AiT)以增强稀疏注意力输入补丁之间的关联性,提升参数效率及关联推理任务性能。AiT利用由多种专门先验组成的可学习显式记忆,结合瓶颈注意力机制,促进多样化局部特征的提取。此外,我们提出一种基于Hopfield能量函数的联想记忆补丁重建新方法。在三种不同规模的AiT上进行的四项图像分类任务的大量实验表明,AiT在显著减少参数和注意力层数量的同时,性能优于Vision Transformers及多种稀疏Transformer。同时,AiT在Sort-of-CLEVR数据集上取得新的最优性能,超越了先前的协调方法。