Permanent faults induced due to imperfections in the manufacturing process of Deep Neural Network (DNN) accelerators are a major concern, as they negatively impact the manufacturing yield of the chip fabrication process. Fault-aware training is the state-of-the-art approach for mitigating such faults. However, it incurs huge retraining overheads, specifically when used for large DNNs trained on complex datasets. To address this issue, we propose a novel Fault-Aware Quantization (FAQ) technique for mitigating the effects of stuck-at permanent faults in the on-chip weight memory of DNN accelerators at a negligible overhead cost compared to fault-aware retraining while offering comparable accuracy results. We propose a lookup table-based algorithm to achieve ultra-low model conversion time. We present extensive evaluation of the proposed approach using five different DNNs, i.e., ResNet-18, VGG11, VGG16, AlexNet and MobileNetV2, and three different datasets, i.e., CIFAR-10, CIFAR-100 and ImageNet. The results demonstrate that FAQ helps in maintaining the baseline accuracy of the DNNs at low and moderate fault rates without involving costly fault-aware training. For example, for ResNet-18 trained on the CIFAR-10 dataset, at 0.04 fault rate FAQ offers (on average) an increase of 76.38% in accuracy. Similarly, for VGG11 trained on the CIFAR-10 dataset, at 0.04 fault rate FAQ offers (on average) an increase of 70.47% in accuracy. The results also show that FAQ incurs negligible overheads, i.e., less than 5% of the time required to run 1 epoch of retraining. We additionally demonstrate the efficacy of our technique when used in conjunction with fault-aware retraining and show that the use of FAQ inside fault-aware retraining enables fast accuracy recovery.
翻译:深度神经网络(DNN)加速器制造工艺中的缺陷导致的永久性故障是主要问题,因为它们会降低芯片制造过程的良率。故障感知训练是缓解此类故障的当前最优方法,但该方法会带来巨大的重训练开销,尤其在处理基于复杂数据集训练的大型DNN时。为解决此问题,我们提出一种新颖的故障感知量化(FAQ)技术,用于缓解DNN加速器片上权重存储器中的固定型永久性故障影响。与故障感知重训练相比,该技术仅需极低的额外开销,同时能提供相近的精度结果。我们提出了一种基于查找表的算法,以实现超低的模型转换时间。我们使用五种不同的DNN(ResNet-18、VGG11、VGG16、AlexNet和MobileNetV2)和三种数据集(CIFAR-10、CIFAR-100和ImageNet)对所提方法进行了全面评估。结果表明,FAQ有助于在低、中故障率下保持DNN的基线精度,且无需进行成本高昂的故障感知训练。例如,对于在CIFAR-10数据集上训练的ResNet-18,在0.04故障率下,FAQ平均提升76.38%的精度。类似地,对于在CIFAR-10数据集上训练的VGG11,在0.04故障率下,FAQ平均提升70.47%的精度。结果还显示,FAQ产生的开销可忽略不计——所需时间不到运行1个重训练轮次的5%。我们还展示了该技术与故障感知重训练结合使用的有效性,并证明在故障感知重训练中引入FAQ能够实现快速精度恢复。