Softmax attention is the principle backbone of foundation models for various artificial intelligence applications, yet its quadratic complexity in sequence length can limit its inference throughput in long-context settings. To address this challenge, alternative architectures such as linear attention, State Space Models (SSMs), and Recurrent Neural Networks (RNNs) have been considered as more efficient alternatives. While connections between these approaches exist, such models are commonly developed in isolation and there is a lack of theoretical understanding of the shared principles underpinning these architectures and their subtle differences, greatly influencing performance and scalability. In this paper, we introduce the Dynamical Systems Framework (DSF), which allows a principled investigation of all these architectures in a common representation. Our framework facilitates rigorous comparisons, providing new insights on the distinctive characteristics of each model class. For instance, we compare linear attention and selective SSMs, detailing their differences and conditions under which both are equivalent. We also provide principled comparisons between softmax attention and other model classes, discussing the theoretical conditions under which softmax attention can be approximated. Additionally, we substantiate these new insights with empirical validations and mathematical arguments. This shows the DSF's potential to guide the systematic development of future more efficient and scalable foundation models.
翻译:Softmax注意力机制是各类人工智能应用中基础模型的核心架构,但其在序列长度上的二次复杂度会限制其在长上下文场景中的推理吞吐量。为应对这一挑战,线性注意力、状态空间模型(SSMs)和循环神经网络(RNNs)等替代架构已被视为更高效的方案。尽管这些方法之间存在关联,但它们通常被孤立地发展,且缺乏对这些架构所共享的基本原理及其细微差异的理论理解,而这些差异极大地影响着模型的性能与可扩展性。本文中,我们提出了动态系统框架(DSF),该框架允许在统一的表示形式中对所有这些架构进行系统性研究。我们的框架促进了严谨的比较,为每类模型的独特特性提供了新的见解。例如,我们比较了线性注意力与选择性SSMs,详细阐述了两者的差异及其等价的条件。我们还对Softmax注意力与其他模型类别进行了原理性比较,讨论了Softmax注意力可被近似的理论条件。此外,我们通过实证验证与数学论证支持了这些新见解。这表明DSF具有指导未来更高效、可扩展基础模型系统化发展的潜力。