Vision Transformers (ViTs) have demonstrated the state-of-the-art performance in various vision-related tasks. The success of ViTs motivates adversaries to perform backdoor attacks on ViTs. Although the vulnerability of traditional CNNs to backdoor attacks is well-known, backdoor attacks on ViTs are seldom-studied. Compared to CNNs capturing pixel-wise local features by convolutions, ViTs extract global context information through patches and attentions. Na\"ively transplanting CNN-specific backdoor attacks to ViTs yields only a low clean data accuracy and a low attack success rate. In this paper, we propose a stealth and practical ViT-specific backdoor attack $TrojViT$. Rather than an area-wise trigger used by CNN-specific backdoor attacks, TrojViT generates a patch-wise trigger designed to build a Trojan composed of some vulnerable bits on the parameters of a ViT stored in DRAM memory through patch salience ranking and attention-target loss. TrojViT further uses minimum-tuned parameter update to reduce the bit number of the Trojan. Once the attacker inserts the Trojan into the ViT model by flipping the vulnerable bits, the ViT model still produces normal inference accuracy with benign inputs. But when the attacker embeds a trigger into an input, the ViT model is forced to classify the input to a predefined target class. We show that flipping only few vulnerable bits identified by TrojViT on a ViT model using the well-known RowHammer can transform the model into a backdoored one. We perform extensive experiments of multiple datasets on various ViT models. TrojViT can classify $99.64\%$ of test images to a target class by flipping $345$ bits on a ViT for ImageNet.
翻译:视觉Transformer(ViT)已在各类视觉任务中展现出最先进的性能。ViT的成功促使攻击者对其发起后门攻击。尽管传统卷积神经网络(CNN)易受后门攻击的弱点已广为人知,但针对ViT的后门攻击却鲜有研究。与通过卷积捕获像素级局部特征的CNN不同,ViT通过图像块和注意力机制提取全局上下文信息。将CNN特有的后门攻击简单移植到ViT上,仅能获得较低的干净数据准确率和攻击成功率。本文提出一种隐蔽且实用的ViT专用后门攻击方法$TrojViT$。与CNN专用后门攻击使用的区域级触发器不同,TrojViT生成一种图像块级触发器,通过图像块显著性排序和注意力目标损失,在存储于DRAM中的ViT参数上构建由若干脆弱比特组成的特洛伊木马。TrojViT进一步采用最小调参更新策略,以减少特洛伊木马的比特数量。一旦攻击者通过翻转这些脆弱比特将特洛伊木马植入ViT模型,该模型对良性输入仍能保持正常推理精度;但当攻击者向输入嵌入触发器时,ViT模型将强制将该输入分类至预设目标类别。我们证明,仅需翻转TrojViT识别的少量脆弱比特(利用著名的RowHammer技术),即可将ViT模型转化为带后门模型。我们在多种ViT模型上对多个数据集进行了大量实验。针对ImageNet数据集,TrojViT仅需翻转ViT模型中的$345$个比特,即可使$99.64\%$的测试图像被分类至目标类别。