Modern deep networks are trained with stochastic gradient descent (SGD) whose key hyperparameters are the number of data considered at each step or batch size $B$, and the step size or learning rate $\eta$. For small $B$ and large $\eta$, SGD corresponds to a stochastic evolution of the parameters, whose noise amplitude is governed by the ''temperature'' $T\equiv \eta/B$. Yet this description is observed to break down for sufficiently large batches $B\geq B^*$, or simplifies to gradient descent (GD) when the temperature is sufficiently small. Understanding where these cross-overs take place remains a central challenge. Here, we resolve these questions for a teacher-student perceptron classification model and show empirically that our key predictions still apply to deep networks. Specifically, we obtain a phase diagram in the $B$-$\eta$ plane that separates three dynamical phases: (i) a noise-dominated SGD governed by temperature, (ii) a large-first-step-dominated SGD and (iii) GD. These different phases also correspond to different regimes of generalization error. Remarkably, our analysis reveals that the batch size $B^*$ separating regimes (i) and (ii) scale with the size $P$ of the training set, with an exponent that characterizes the hardness of the classification problem.
翻译:现代深度网络通过随机梯度下降(SGD)进行训练,其关键超参数为每次迭代处理的数据量(即批次大小 $B$)和步长(即学习率 $\eta$)。当 $B$ 较小且 $\eta$ 较大时,SGD对应于参数的随机演化过程,其噪声幅度由"温度" $T\equiv \eta/B$ 主导。然而,当批次足够大(即 $B\geq B^*$)时,该描述被观察到失效;而当温度足够小时,其行为简化为梯度下降(GD)。理解这些机制转变发生的临界点仍是核心挑战。本文以教师-学生感知机分类模型为研究对象,解析了上述问题,并通过实证表明我们的关键预测仍适用于深度网络。具体而言,我们在 $B$-$\eta$ 相图中分离出三种动力学相态:(i) 由温度主导的噪声型SGD;(ii) 首步主导型SGD;(iii) 梯度下降(GD)。这些不同相态对应差异化的泛化误差机制。值得关注的是,分析表明区分相态(i)与(ii)的临界批次大小 $B^*$ 与训练集规模 $P$ 成比例缩放,其指数特性反映了分类问题的难易程度。