We study the training dynamics of shallow neural networks, investigating the conditions under which a limited number of large batch gradient descent steps can facilitate feature learning beyond the kernel regime. We compare the influence of batch size and that of multiple (but finitely many) steps. Our analysis of a single-step process reveals that while a batch size of $n = O(d)$ enables feature learning, it is only adequate for learning a single direction, or a single-index model. In contrast, $n = O(d^2)$ is essential for learning multiple directions and specialization. Moreover, we demonstrate that ``hard'' directions, which lack the first $\ell$ Hermite coefficients, remain unobserved and require a batch size of $n = O(d^\ell)$ for being captured by gradient descent. Upon iterating a few steps, the scenario changes: a batch-size of $n = O(d)$ is enough to learn new target directions spanning the subspace linearly connected in the Hermite basis to the previously learned directions, thereby a staircase property. Our analysis utilizes a blend of techniques related to concentration, projection-based conditioning, and Gaussian equivalence that are of independent interest. By determining the conditions necessary for learning and specialization, our results highlight the interaction between batch size and number of iterations, and lead to a hierarchical depiction where learning performance exhibits a stairway to accuracy over time and batch size, shedding new light on feature learning in neural networks.
翻译:我们研究了浅层神经网络的训练动态,探究了在何种条件下,有限次的大批量梯度下降步骤能够促进超越核机制的特征学习。我们比较了批量大小与有限步数的影响。对单步过程的分析表明:当批量大小 $n = O(d)$ 时虽能实现特征学习,但仅适用于学习单一方向(即单指标模型);而 $n = O(d^2)$ 对于学习多方向及其特化至关重要。此外,我们证明"硬"方向(缺乏前 $\ell$ 个埃尔米特系数)在梯度下降中无法被观测,需要批量大小达到 $n = O(d^\ell)$ 才能捕获。当迭代多步后,情况发生变化:$n = O(d)$ 的批量足以学习新的目标方向——这些方向在埃尔米特基上与先前学习的方向线性连接,形成阶梯特性。我们的分析融合了浓度不等式、基于投影的条件化、高斯等价性等独立课题的技术。通过确定学习与特化的必要条件,结果揭示了批量大小与迭代次数的交互作用,并构建了层次化图景:学习性能随时间和批量大小呈现"准确度阶梯",为神经网络中的特征学习提供了新见解。