We investigate the ability of deep neural networks to identify the support of the target function. Our findings reveal that mini-batch SGD effectively learns the support in the first layer of the network by shrinking to zero the weights associated with irrelevant components of input. In contrast, we demonstrate that while vanilla GD also approximates the target function, it requires an explicit regularization term to learn the support in the first layer. We prove that this property of mini-batch SGD is due to a second-order implicit regularization effect which is proportional to $\eta / b$ (step size / batch size). Our results are not only another proof that implicit regularization has a significant impact on training optimization dynamics but they also shed light on the structure of the features that are learned by the network. Additionally, they suggest that smaller batches enhance feature interpretability and reduce dependency on initialization.
翻译:本研究探讨了深度神经网络识别目标函数支撑集的能力。我们的研究结果表明,小批量SGD通过将输入无关分量的权重收缩至零,能有效在网络第一层学习支撑集。相比之下,我们证明虽然标准GD也能逼近目标函数,但其需要在第一层引入显式正则化项才能学习支撑集。我们证实小批量SGD的这一特性源于二阶隐式正则化效应,该效应与$\eta / b$(步长/批次大小)成正比。这些结果不仅再次证明了隐式正则化对训练优化动态具有重要影响,还揭示了网络所学特征的结构特性。此外,研究还表明较小批次规模能增强特征可解释性并降低对初始化的依赖性。