We attribute grokking, the phenomenon where generalization is much delayed after memorization, to compression. To do so, we define linear mapping number (LMN) to measure network complexity, which is a generalized version of linear region number for ReLU networks. LMN can nicely characterize neural network compression before generalization. Although the $L_2$ norm has been a popular choice for characterizing model complexity, we argue in favor of LMN for a number of reasons: (1) LMN can be naturally interpreted as information/computation, while $L_2$ cannot. (2) In the compression phase, LMN has linear relations with test losses, while $L_2$ is correlated with test losses in a complicated nonlinear way. (3) LMN also reveals an intriguing phenomenon of the XOR network switching between two generalization solutions, while $L_2$ does not. Besides explaining grokking, we argue that LMN is a promising candidate as the neural network version of the Kolmogorov complexity since it explicitly considers local or conditioned linear computations aligned with the nature of modern artificial neural networks.
翻译:我们将突现(grokking)——即泛化远滞后于记忆的现象——归因于压缩。为此,我们定义了线性映射数(LMN)来衡量网络复杂性,这是ReLU网络线性区域数的广义版本。LMN能很好地刻画神经网络在泛化前的压缩过程。尽管L2范数长期以来是描述模型复杂性的常用选择,但我们基于以下理由主张采用LMN:(1)LMN可自然地解释为信息/计算量,而L2范数则不能;(2)在压缩阶段,LMN与测试损失呈线性关系,而L2范数与测试损失却以复杂的非线性方式相关;(3)LMN还揭示了XOR网络在两个泛化解之间切换的有趣现象,而L2范数则无法展现。除了解释突现现象之外,我们认为LMN是科尔莫戈罗夫复杂性的神经网络版的有力候选者,因为它显式地考虑了与现代人工神经网络性质相符的局部或条件线性计算。