Linear attention is an efficient attention mechanism that has recently emerged as a promising alternative to conventional softmax attention. With its ability to process tokens in linear computational complexities, linear attention, in theory, can handle sequences of unlimited length without sacrificing speed, i.e., maintaining a constant training speed for various sequence lengths with a fixed memory consumption. However, due to the issue with cumulative summation (cumsum), current linear attention algorithms cannot demonstrate their theoretical advantage in a causal setting. In this paper, we present Lightning Attention-2, the first linear attention implementation that enables linear attention to realize its theoretical computational benefits. To achieve this, we leverage the thought of tiling, separately handling the intra-block and inter-block components in linear attention calculation. Specifically, we utilize the conventional attention computation mechanism for the intra-blocks and apply linear attention kernel tricks for the inter-blocks. A tiling technique is adopted through both forward and backward procedures to take full advantage of the GPU hardware. We implement our algorithm in Triton to make it IO-aware and hardware-friendly. Various experiments are conducted on different model sizes and sequence lengths. Lightning Attention-2 retains consistent training and inference speed regardless of input sequence length and is significantly faster than other attention mechanisms. The source code is available at https://github.com/OpenNLPLab/lightning-attention.
翻译:线性注意力是一种高效的注意力机制,近期已成为传统softmax注意力的有前景的替代方案。凭借以线性计算复杂度处理token的能力,线性注意力理论上能够处理无限长度的序列而不损失速度——即在固定内存消耗下,对不同序列长度保持恒定的训练速度。然而,由于累积求和(cumsum)问题,现有的线性注意力算法在因果设置下无法展现其理论优势。本文提出Lightning Attention-2,这是首个使线性注意力实现其理论计算优势的线性注意力实现方案。为实现这一目标,我们利用分块思想,在线性注意力计算中分别处理块内和块间组件。具体而言,我们对块内组件采用传统注意力计算机制,对块间组件应用线性注意力核技巧。通过前向和反向传播过程采用分块技术,以充分利用GPU硬件。我们使用Triton实现算法,使其具备IO感知性和硬件友好性。在不同模型规模和序列长度上进行了多种实验。Lightning Attention-2无论输入序列长度如何,均能保持一致稳定的训练和推理速度,且显著快于其他注意力机制。源代码已开源,见https://github.com/OpenNLPLab/lightning-attention。