In deep learning theory, the covariance matrix of the representations serves as a proxy to examine the network's trainability. Motivated by the success of Transformers, we study the covariance matrix of a modified Softmax-based attention model with skip connections in the proportional limit of infinite-depth-and-width. We show that at initialization the limiting distribution can be described by a stochastic differential equation (SDE) indexed by the depth-to-width ratio. To achieve a well-defined stochastic limit, the Transformer's attention mechanism is modified by centering the Softmax output at identity, and scaling the Softmax logits by a width-dependent temperature parameter. We examine the stability of the network through the corresponding SDE, showing how the scale of both the drift and diffusion can be elegantly controlled with the aid of residual connections. The existence of a stable SDE implies that the covariance structure is well-behaved, even for very large depth and width, thus preventing the notorious issues of rank degeneracy in deep attention models. Finally, we show, through simulations, that the SDE provides a surprisingly good description of the corresponding finite-size model. We coin the name shaped Transformer for these architectural modifications.
翻译:在深度学习理论中,表征的协方差矩阵作为考察网络可训练性的代理指标。受Transformer成功经验的启发,我们研究了一种基于Softmax的修正注意力模型(含跳跃连接)在深度与宽度成比例增长的无限极限下的协方差矩阵。我们证明,在初始化阶段,其极限分布可由以深度-宽度比为索引的随机微分方程(SDE)描述。为实现良定义的随机极限,我们通过将Softmax输出中心化至单位矩阵,并采用宽度相关温度参数缩放Softmax对数几率来修正Transformer的注意力机制。我们通过对应的SDE考察网络稳定性,展示如何借助残差连接优雅地控制漂移项与扩散项的尺度。稳定SDE的存在性表明协方差结构在极大深度与宽度下仍保持良好性质,从而规避深度注意力模型中臭名昭著的秩退化问题。最后,我们通过仿真证明该SDE能出奇准确地描述对应有限尺寸模型的行为。我们将这种架构改进命名为形状Transformer。