We study the evolution of hidden-weight spectra in wide neural networks trained by (stochastic) gradient descent. We develop a two-level dynamical mean-field theory (DMFT) that jointly tracks bulk and outlier spectral dynamics for spiked ensembles whose spike directions remain statistically dependent on the random bulk. We apply this framework to two settings: (1) infinite-width nonlinear networks in mean-field/$μ$P scaling and (2) deep linear networks in the proportional high-dimensional limit, where width, input dimension, and sample size diverge with fixed ratios. Our theory predicts how outliers evolve with training time, width, output scale, and initialization variance. In deep linear networks, $μ$P yields width-consistent outlier dynamics and hyperparameter transfer, including width-stable growth of the leading NTK mode toward the edge of stability (EoS). In contrast, NTK parameterization exhibits strongly width-dependent outlier dynamics, despite converging to a stable large-width limit. We show that this bulk+outlier picture is descriptive of simple tasks with small output channels, but that tasks involving large numbers of outputs (ImageNet classification or GPT language modeling) are better described by a restructuring of the spectral bulk. We develop a toy model with extensive output channels that recapitulates this phenomenon and show that edge of the spectrum still converges for sufficiently wide networks.
翻译:我们研究了由(随机)梯度下降训练的宽神经网络中隐藏权重的谱演化过程。我们发展了一种双层动态平均场理论(DMFT),该理论能够联合追踪尖峰系综的体谱与离群谱动力学,其中尖峰方向仍与随机体谱保持统计依赖关系。我们将此框架应用于两种场景:(1)平均场/μP缩放下的无穷宽非线性网络;(2)比例高维极限下的深度线性网络,其中宽度、输入维度和样本量以固定比率发散。我们的理论预测了离群值如何随训练时间、宽度、输出尺度和初始化方差演化。在深度线性网络中,μP产生宽度一致的离群动力学和超参数迁移,包括主导NTK模式向稳定性边界(EoS)的宽度稳定增长。相比之下,NTK参数化表现出强烈的宽度依赖性离群动力学,尽管收敛到稳定的宽宽度极限。我们证明这种体+离群图景能描述小输出通道的简单任务,但涉及大量输出(ImageNet分类或GPT语言建模)的任务更适合用谱体重构来描述。我们开发了一个具有广泛输出通道的玩具模型来重现这一现象,并证明对于足够宽的网络,谱边界仍然收敛。