This paper focuses on predicting the occurrence of grokking in neural networks, a phenomenon in which perfect generalization emerges long after signs of overfitting or memorization are observed. It has been reported that grokking can only be observed with certain hyper-parameters. This makes it critical to identify the parameters that lead to grokking. However, since grokking occurs after a large number of epochs, searching for the hyper-parameters that lead to it is time-consuming. In this paper, we propose a low-cost method to predict grokking without training for a large number of epochs. In essence, by studying the learning curve of the first few epochs, we show that one can predict whether grokking will occur later on. Specifically, if certain oscillations occur in the early epochs, one can expect grokking to occur if the model is trained for a much longer period of time. We propose using the spectral signature of a learning curve derived by applying the Fourier transform to quantify the amplitude of low-frequency components to detect the presence of such oscillations. We also present additional experiments aimed at explaining the cause of these oscillations and characterizing the loss landscape.
翻译:本文聚焦于预测神经网络中“顿悟”(grokking)现象的发生,即完美泛化能力在观察到过拟合或记忆化迹象很久之后才突然涌现的现象。已有研究表明,“顿悟”仅能在特定超参数设置下被观察到。因此,识别导致该现象的参数至关重要。然而,由于“顿悟”需要经过大量训练轮次才会出现,搜索引发该现象的超参数极为耗时。本文提出了一种低成本方法,无需大量训练轮次即可预测“顿悟”的发生。本质而言,通过研究前几轮训练的学习曲线,我们证明可以预测后续是否会出现“顿悟”。具体而言,若在早期训练轮次中观察到特定振荡,则预期模型在更长时间训练后会出现“顿悟”。我们提出利用傅里叶变换提取学习曲线的频谱特征,通过量化低频分量的振幅来检测这类振荡的存在。此外,我们通过补充实验解释这些振荡的成因并对损失景观进行刻画。