We propose a new class of linear Transformers called FourierLearner-Transformers (FLTs), which incorporate a wide range of relative positional encoding mechanisms (RPEs). These include regular RPE techniques applied for nongeometric data, as well as novel RPEs operating on the sequences of tokens embedded in higher-dimensional Euclidean spaces (e.g. point clouds). FLTs construct the optimal RPE mechanism implicitly by learning its spectral representation. As opposed to other architectures combining efficient low-rank linear attention with RPEs, FLTs remain practical in terms of their memory usage and do not require additional assumptions about the structure of the RPE-mask. FLTs allow also for applying certain structural inductive bias techniques to specify masking strategies, e.g. they provide a way to learn the so-called local RPEs introduced in this paper and providing accuracy gains as compared with several other linear Transformers for language modeling. We also thoroughly tested FLTs on other data modalities and tasks, such as: image classification and 3D molecular modeling. For 3D-data FLTs are, to the best of our knowledge, the first Transformers architectures providing RPE-enhanced linear attention.
翻译:我们提出一类新型线性Transformer架构——傅里叶学习变换器(FLTs),该架构融合了广泛的相对位置编码(RPE)机制。这些机制包括应用于非几何数据的传统RPE技术,以及针对嵌入高维欧几里得空间(如点云)的令牌序列的新型RPE。FLTs通过隐式学习其频谱表示来构建最优RPE机制。与结合高效低秩线性注意力与RPE的其他架构不同,FLTs在内存使用方面保持实用性,且无需对RPE掩码结构引入额外假设。FLTs还可运用特定结构归纳偏置技术来指定掩码策略,例如本文提出的局部RPE可通过该方法学习,相比其他线性Transformer在语言建模中实现精度提升。我们还在图像分类与3D分子建模等数据模态与任务上对FLTs进行了充分测试。据我们所知,对于3D数据,FLTs是首个提供RPE增强线性注意力的Transformer架构。