Vision Transformers (ViTs) that leverage self-attention mechanism have shown superior performance on many classical vision tasks compared to convolutional neural networks (CNNs) and gain increasing popularity recently. Existing ViTs works mainly optimize performance and accuracy, but ViTs reliability issues induced by soft errors in large-scale VLSI designs have generally been overlooked. In this work, we mainly study the reliability of ViTs and investigate the vulnerability from different architecture granularities ranging from models, layers, modules, and patches for the first time. The investigation reveals that ViTs with the self-attention mechanism are generally more resilient on linear computing including general matrix-matrix multiplication (GEMM) and full connection (FC) and show a relatively even vulnerability distribution across the patches. ViTs involve more fragile non-linear computing such as softmax and GELU compared to typical CNNs. With the above observations, we propose a lightweight block-wise algorithm-based fault tolerance (LB-ABFT) approach to protect the linear computing implemented with distinct sizes of GEMM and apply a range-based protection scheme to mitigate soft errors in non-linear computing. According to our experiments, the proposed fault-tolerant approaches enhance ViTs accuracy significantly with minor computing overhead in presence of various soft errors.
翻译:视觉Transformer(ViTs)利用自注意力机制在众多经典视觉任务中展现出优于卷积神经网络(CNNs)的性能,近年来日益流行。现有ViTs工作主要优化性能和精度,但由大规模VLSI设计中的软错误引发的ViTs可靠性问题普遍被忽视。本文主要研究ViTs的可靠性,并首次从模型、层、模块和补丁等不同架构粒度探究其脆弱性。研究发现,具有自注意力机制的ViTs在一般矩阵乘法(GEMM)和全连接(FC)等线性计算上普遍更具弹性,且补丁间的脆弱性分布相对均匀。与典型CNN相比,ViTs涉及更多脆弱的非线性计算,如softmax和GELU。基于上述观察,我们提出一种轻量级的基于块算法容错方法(LB-ABFT),以保护以不同规模GEMM实现的线性计算,并应用基于范围的保护方案来缓解非线性计算中的软错误。实验表明,在各种软错误存在的情况下,所提出的容错方法在轻微计算开销下显著提升了ViTs的精度。