The Edge-of-Chaos (EoC) theory developed for the random initialization of deep networks allows more efficient training by both preserving information in the initial outputs of the network and minimising exploding or vanishing gradients through characterisation of the intermediate layers as Gaussian processes. This EoC theory provides formulae for the choice of the initialisation distribution variances of the weights and biases. For activations which are approximately linear around the origin, the EoC theory typically encourages the Gaussian process variance to converge towards zero with increasing depth. Here we consider the less studied setting of highly sparsity inducing activations where a large region of values near the origin are set to zero. In this setting we prove a new phenomenon whereby initialisations leading to larger fixed Gaussian processes are beneficial to training stability. This theory informs a new, yet simple, initialisation strategy that allows training DNNs and CNNs with as large as 90\% sparsity in the hidden layers.
翻译:针对深度网络随机初始化开发的边缘混沌(Edge-of-Chaos, EoC)理论,通过将中间层表征为高斯过程以保留网络初始输出信息并最小化梯度消失/爆炸,从而支持更高效的训练。该理论给出了权重与偏置初始化分布方差的选择公式。对于在原点附近近似线性的激活函数,EoC理论通常促使高斯过程方差随深度增加趋于零。本文考虑较少研究的强稀疏性诱导激活函数场景——在此类激活函数中,原点附近大范围区域的值被置为零。我们在此场景下证明了一个新现象:导致较大固定高斯过程的初始化更有利于训练稳定性。该理论指导了一种新颖且简单的初始化策略,使得深度神经网络和卷积神经网络在隐藏层稀疏度高达90%时仍能保持训练稳定性。