A central theme of the modern machine learning paradigm is that larger neural networks achieve better performance on a variety of metrics. Theoretical analyses of these overparameterized models have recently centered around studying very wide neural networks. In this tutorial, we provide a nonrigorous but illustrative derivation of the following fact: in order to train wide networks effectively, there is only one degree of freedom in choosing hyperparameters such as the learning rate and the size of the initial weights. This degree of freedom controls the richness of training behavior: at minimum, the wide network trains lazily like a kernel machine, and at maximum, it exhibits feature learning in the so-called $\mu$P regime. In this paper, we explain this richness scale, synthesize recent research results into a coherent whole, offer new perspectives and intuitions, and provide empirical evidence supporting our claims. In doing so, we hope to encourage further study of the richness scale, as it may be key to developing a scientific theory of feature learning in practical deep neural networks.
翻译:现代机器学习范式的一个核心主题是:更大的神经网络能够在多种指标上实现更优的性能。针对这些过参数化模型的理论分析近期集中在研究极宽神经网络上。在本教程中,我们提供了一种非严格但具有启发性的推导,阐释以下事实:为有效训练宽网络,超参数(如学习率和初始权重尺度)仅存在一个自由度。该自由度控制着训练行为的丰富程度:在最小值时,宽网络的训练呈现类似核机器的懒惰行为;在最大值时,则表现出所谓$μ$P机制下的特征学习。本文阐释了这一丰富度标度,将近期研究成果综合为统一框架,提出新视角与直观认识,并提供支持我们结论的实证证据。通过此举,我们期望推动对丰富度标度的进一步研究——这或许正是为实际深度神经网络中的特征学习建立科学理论的关键所在。