Backpropagation (BP) is the dominant and most successful method for training parameters of deep neural network models. However, BP relies on two computationally distinct phases, does not provide a satisfactory explanation of biological learning, and can be challenging to apply for training of networks with discontinuities or noisy node dynamics. By comparison, node perturbation (NP) proposes learning by the injection of noise into network activations, and subsequent measurement of the induced loss change. NP relies on two forward (inference) passes, does not make use of network derivatives, and has been proposed as a model for learning in biological systems. However, standard NP is highly data inefficient and unstable due to its unguided noise-based search process. In this work, we investigate different formulations of NP and relate it to the concept of directional derivatives as well as combining it with a decorrelating mechanism for layer-wise inputs. We find that a closer alignment with directional derivatives together with input decorrelation at every layer significantly enhances performance of NP learning with significant improvements in parameter convergence and much higher performance on the test data, approaching that of BP. Furthermore, our novel formulation allows for application to noisy systems in which the noise process itself is inaccessible.
翻译:反向传播(BP)是训练深度神经网络参数的主导且最成功的方法。然而,BP依赖于两个计算上截然不同的阶段,无法为生物学习提供满意的解释,并且在训练具有不连续性或噪声节点动力学的网络时可能具有挑战性。相比之下,节点扰动(NP)通过向网络激活中注入噪声并随后测量由此产生的损失变化来提出学习方法。NP依赖于两次前向(推理)传递,不利用网络导数,并已被提议作为生物系统中学习的模型。然而,标准NP由于其无引导的基于噪声的搜索过程而数据效率低下且不稳定。在这项工作中,我们研究了NP的不同形式,并将其与方向导数的概念联系起来,同时将其与层间输入的去相关机制相结合。我们发现,与方向导数的更紧密对齐以及每层输入的去相关显著提升了NP学习的性能,在参数收敛方面有显著改进,并在测试数据上达到了接近BP的更高性能。此外,我们的新颖公式允许应用于噪声过程本身不可访问的噪声系统。