By classifying infinite-width neural networks and identifying the *optimal* limit, Tensor Programs IV and V demonstrated a universal way, called $\mu$P, for *widthwise hyperparameter transfer*, i.e., predicting optimal hyperparameters of wide neural networks from narrow ones. Here we investigate the analogous classification for *depthwise parametrizations* of deep residual networks (resnets). We classify depthwise parametrizations of block multiplier and learning rate by their infinite-width-then-depth limits. In resnets where each block has only one layer, we identify a unique optimal parametrization, called Depth-$\mu$P that extends $\mu$P and show empirically it admits depthwise hyperparameter transfer. We identify *feature diversity* as a crucial factor in deep networks, and Depth-$\mu$P can be characterized as maximizing both feature learning and feature diversity. Exploiting this, we find that absolute value, among all homogeneous nonlinearities, maximizes feature diversity and indeed empirically leads to significantly better performance. However, if each block is deeper (such as modern transformers), then we find fundamental limitations in all possible infinite-depth limits of such parametrizations, which we illustrate both theoretically and empirically on simple networks as well as Megatron transformer trained on Common Crawl.
翻译:通过对无限宽度神经网络进行分类并识别*最优*极限,张量程序IV和V提出了一种称为$\mu$P的通用方法,用于*宽度方向超参数迁移*,即从窄网络预测宽网络的最优超参数。本文研究了深层残差网络(resnets)的*深度方向参数化*的类似分类。我们根据无限宽度后取深度极限,对残差网络中块乘子与学习率的深度方向参数化进行分类。在每块仅含一层的残差网络中,我们识别出唯一的最优参数化——称为Depth-$\mu$P(扩展了$\mu$P),并通过实验证明其支持深度方向超参数迁移。我们将*特征多样性*认定为深度网络的关键因素,而Depth-$\mu$P可被表征为同时最大化特征学习与特征多样性。利用这一特性,我们发现所有齐次非线性函数中,绝对值函数能最大化特征多样性,并且实验上确实显著提升性能。然而,若每个块更深(如现代Transformer),则此类参数化的所有无限深度极限均存在根本性局限——我们在简单网络以及基于Common Crawl训练的Megatron Transformer上,均从理论与实验两方面阐明了这一现象。