Deep neural networks (DNN) with a huge number of adjustable parameters remain largely black boxes. To shed light on the hidden layers of DNN, we study supervised learning by a DNN of width $N$ and depth $L$ consisting of $NL$ perceptrons with $c$ inputs by a statistical mechanics approach called the teacher-student setting. We consider an ensemble of student machines that exactly reproduce $M$ sets of $N$ dimensional input/output relations provided by a teacher machine. We show that the problem becomes exactly solvable in what we call as 'dense limit': $N \gg c \gg 1$ and $M \gg 1$ with fixed $\alpha=M/c$ using the replica method developed in (H. Yoshino, (2020)). We also study the model numerically performing simple greedy MC simulations. Simulations reveal that learning by the DNN is quite heterogeneous in the network space: configurations of the teacher and the student machines are more correlated within the layers closer to the input/output boundaries while the central region remains much less correlated due to the over-parametrization in qualitative agreement with the theoretical prediction. We evaluate the generalization-error of the DNN with various depth $L$ both theoretically and numerically. Remarkably both the theory and simulation suggest generalization-ability of the student machines, which are only weakly correlated with the teacher in the center, does not vanish even in the deep limit $L \gg 1$ where the system becomes heavily over-parametrized. We also consider the impact of effective dimension $D(\leq N)$ of data by incorporating the hidden manifold model (S. Goldt et. al., (2020)) into our model. The theory implies that the loop corrections to the dense limit become enhanced by either decreasing the width $N$ or decreasing the effective dimension $D$ of the data. Simulation suggests both lead to significant improvements in generalization-ability.
翻译:拥有大量可调参数的深度神经网络(DNN)在很大程度上仍是黑箱。为揭示DNN隐藏层的机制,我们通过统计力学方法(即师生设定)研究了一个由$c$个输入、$NL$个感知器构成的宽度为$N$、深度为$L$的DNN的监督学习过程。我们考虑一个学生机器集成,能够精确再现教师机器提供的$M$组$N$维输入/输出关系。我们证明,在所谓的"稠密极限"下问题可精确求解:即$N \gg c \gg 1$且$M \gg 1$时,固定参数$\alpha=M/c$,采用(Yoshino, 2020)中发展的复制方法。我们还通过简单贪心MC模拟对模型进行数值研究。模拟揭示DNN学习在网络空间中呈现显著异质性:教师与学生机器的配置在靠近输入/输出边界的层中相关度更高,而由于过度参数化,中间区域的相关度显著降低,这一结果与理论预测定性一致。我们从理论和数值两方面评估不同深度$L$下DNN的泛化误差。值得注意的是,理论和模拟均表明,在深度极限$L \gg 1$(系统严重过参数化)下,中心区域与教师仅弱相关的学生机器的泛化能力并未消失。我们还通过将隐流形模型(Goldt等,2020)纳入我们的模型,考虑了数据有效维度$D(\leq N)$的影响。理论表明,减小宽度$N$或数据有效维度$D$均会增强稠密极限下的圈修正。模拟表明两者均能显著提升泛化能力。