The remarkable performance of overparameterized deep neural networks (DNNs) must arise from an interplay between network architecture, training algorithms, and structure in the data. To disentangle these three components, we apply a Bayesian picture, based on the functions expressed by a DNN, to supervised learning. The prior over functions is determined by the network, and is varied by exploiting a transition between ordered and chaotic regimes. For Boolean function classification, we approximate the likelihood using the error spectrum of functions on data. When combined with the prior, this accurately predicts the posterior, measured for DNNs trained with stochastic gradient descent. This analysis reveals that structured data, combined with an intrinsic Occam's razor-like inductive bias towards (Kolmogorov) simple functions that is strong enough to counteract the exponential growth of the number of functions with complexity, is a key to the success of DNNs.
翻译:过参数化深度神经网络(DNN)的卓越性能必然源于网络架构、训练算法与数据结构之间的相互作用。为厘清这三类要素,我们基于深度神经网络所表达的函数,将贝叶斯视角应用于监督学习。函数先验由网络决定,并通过利用有序与混沌态之间的相变进行调节。针对布尔函数分类问题,我们采用函数在数据上的误差谱近似似然函数。该近似与先验结合后,能够精确预测随机梯度下降训练DNN时所测得的后验分布。分析表明:结构化数据与内嵌的、对(柯尔莫哥洛夫)简单函数的奥卡姆剃刀式归纳偏置相结合——该偏置强度足以抵消函数数量随复杂度增长的指数级膨胀——是深度神经网络成功的关键所在。