Deep Learning, and in particular, Deep Neural Network (DNN) is nowadays widely used in many scenarios, including safety-critical applications such as autonomous driving. In this context, besides energy efficiency and performance, reliability plays a crucial role since a system failure can jeopardize human life. As with any other device, the reliability of hardware architectures running DNNs has to be evaluated, usually through costly fault injection campaigns. This paper explores the approximation and fault resiliency of DNN accelerators. We propose to use approximate (AxC) arithmetic circuits to agilely emulate errors in hardware without performing fault injection on the DNN. To allow fast evaluation of AxC DNN, we developed an efficient GPU-based simulation framework. Further, we propose a fine-grain analysis of fault resiliency by examining fault propagation and masking in networks
翻译:深度学习,特别是深度神经网络(DNN),如今广泛应用于众多场景,包括自动驾驶等安全关键型应用。在此背景下,除能效和性能外,可靠性也发挥着至关重要的作用,因为系统故障可能危及人类生命。与其他任何设备类似,运行DNN的硬件架构的可靠性必须通过评估——通常借助代价高昂的故障注入行。本文探究了DNN加速器的近似计算与容错性。我们提出使用近似(AxC)算术电路,在不进行DNN故障注入的情况下,敏捷地模拟硬件错误。为快速评估AxC DNN,我们开发了一个高效的基于GPU的仿真框架。此外,通过分析网络中的故障传播与屏蔽,我们提出了一种细粒度的容错性分析方法。