Deep learning needs high-precision handling of forwarding signals, backpropagating errors, and updating weights. This is inherently required by the learning algorithm since the gradient descent learning rule relies on the chain product of partial derivatives. However, it is challenging to implement deep learning in hardware systems that use noisy analog memristors as artificial synapses, as well as not being biologically plausible. Memristor-based implementations generally result in an excessive cost of neuronal circuits and stringent demands for idealized synaptic devices. Here, we demonstrate that the requirement for high precision is not necessary and that more efficient deep learning can be achieved when this requirement is lifted. We propose a binary stochastic learning algorithm that modifies all elementary neural network operations, by introducing (i) stochastic binarization of both the forwarding signals and the activation function derivatives, (ii) signed binarization of the backpropagating errors, and (iii) step-wised weight updates. Through an extensive hybrid approach of software simulation and hardware experiments, we find that binary stochastic deep learning systems can provide better performance than the software-based benchmarks using the high-precision learning algorithm. Also, the binary stochastic algorithm strongly simplifies the neural network operations in hardware, resulting in an improvement of the energy efficiency for the multiply-and-accumulate operations by more than three orders of magnitudes.
翻译:深度学习需要高精度处理前向信号、反向传播误差及更新权重。这一需求本质源于学习算法本身:梯度下降的学习规则依赖于偏导数的链式乘积。然而,在采用噪声模拟忆阻器作为人工突触的硬件系统中实现深度学习极具挑战性,且缺乏生物学合理性。基于忆阻器的实现通常会导致神经元电路成本过高,并对理想化的突触器件提出严苛要求。本文证明,高精度的要求并非必要,取消这一要求后可以实现更高效的深度学习。我们提出一种二值随机学习算法,通过引入以下创新实现对所有基础神经网络运算的改进:(i) 前向信号与激活函数导数的随机二值化,(ii) 反向传播误差的符号二值化,以及(iii) 阶梯式权重更新。通过软件仿真与硬件实验相结合的混合方法,我们发现基于二值随机学习的深度学习系统能够比使用高精度学习算法的软件基准系统取得更优性能。此外,该二值随机算法显著简化了硬件中的神经网络运算,使乘累加运算的能效提升超过三个数量级。