In this paper we provide, to the best of our knowledge, the first comprehensive approach for incorporating various masking mechanisms into Transformers architectures in a scalable way. We show that recent results on linear causal attention (Choromanski et al., 2021) and log-linear RPE-attention (Luo et al., 2021) are special cases of this general mechanism. However by casting the problem as a topological (graph-based) modulation of unmasked attention, we obtain several results unknown before, including efficient d-dimensional RPE-masking and graph-kernel masking. We leverage many mathematical techniques ranging from spectral analysis through dynamic programming and random walks to new algorithms for solving Markov processes on graphs. We provide a corresponding empirical evaluation.
翻译:在本文中,我们首次提出了一种可扩展地将多种掩码机制集成到Transformer架构中的综合方法。我们证明,近期关于线性因果注意力(Choromanski等人,2021)和对数线性RPE注意力(Luo等人,2021)的研究成果是该通用机制的特例。然而,通过将问题重新表述为对非掩码注意力的拓扑(基于图)调制,我们获得了多项此前未知的结果,包括高效的d维RPE掩码和图核掩码。我们运用了多种数学技术,涵盖从谱分析、动态规划、随机游走到解决图上马尔可夫过程的新算法。我们还提供了相应的实验评估。