Large language models based on the Transformer architecture have demonstrated impressive capabilities to learn in context. However, existing theoretical studies on how this phenomenon arises are limited to the dynamics of a single layer of attention trained on linear regression tasks. In this paper, we study the optimization of a Transformer consisting of a fully connected layer followed by a linear attention layer. The MLP acts as a common nonlinear representation or feature map, greatly enhancing the power of in-context learning. We prove in the mean-field and two-timescale limit that the infinite-dimensional loss landscape for the distribution of parameters, while highly nonconvex, becomes quite benign. We also analyze the second-order stability of mean-field dynamics and show that Wasserstein gradient flow almost always avoids saddle points. Furthermore, we establish novel methods for obtaining concrete improvement rates both away from and near critical points. This represents the first saddle point analysis of mean-field dynamics in general and the techniques are of independent interest.
翻译:基于Transformer架构的大型语言模型展示了在上下文中学习的卓越能力。然而,现有关于这一现象如何产生的理论研究仅限于在线性回归任务上训练的单层注意力动力学。本文研究由全连接层后接线性注意力层组成的Transformer的优化问题。MLP作为通用的非线性表示或特征映射,显著增强了上下文学习的能力。我们在均场和双时间尺度极限下证明,参数分布的无限维损失景观虽高度非凸,但变得相当良性。我们还分析了均场动力学的二阶稳定性,表明Wasserstein梯度流几乎总能避免鞍点。此外,我们建立了在远离和接近临界点时获得具体改进率的新方法。这代表了均场动力学鞍点分析的首次一般性研究,其技术具有独立价值。