To successfully launch backdoor attacks, injected data needs to be correctly labeled; otherwise, they can be easily detected by even basic data filters. Hence, the concept of clean-label attacks was introduced, which is more dangerous as it doesn't require changing the labels of injected data. To the best of our knowledge, the existing clean-label backdoor attacks largely relies on an understanding of the entire training set or a portion of it. However, in practice, it is very difficult for attackers to have it because of training datasets often collected from multiple independent sources. Unlike all current clean-label attacks, we propose a novel clean label method called 'Poison Dart Frog'. Poison Dart Frog does not require access to any training data; it only necessitates knowledge of the target class for the attack, such as 'frog'. On CIFAR10, Tiny-ImageNet, and TSRD, with a mere 0.1\%, 0.025\%, and 0.4\% poisoning rate of the training set size, respectively, Poison Dart Frog achieves a high Attack Success Rate compared to LC, HTBA, BadNets, and Blend. Furthermore, compared to the state-of-the-art attack, NARCISSUS, Poison Dart Frog achieves similar attack success rates without any training data. Finally, we demonstrate that four typical backdoor defense algorithms struggle to counter Poison Dart Frog.
翻译:为成功发起后门攻击,注入数据必须正确标注;否则,即便是基础数据过滤器也能轻易将其检测出来。因此,无需改变注入数据标签的干净标签攻击概念应运而生——这种攻击更具危险性。据我们所知,现有干净标签后门攻击大多依赖对整个训练集或其中部分数据集的认知。然而在实际中,攻击者很难掌握完整训练数据,因为这些数据集通常来自多个独立来源。与现有所有干净标签攻击不同,我们提出名为"毒箭蛙"的新型干净标签方法。该方法无需访问任何训练数据,仅需知道攻击目标类别(如"青蛙")。在CIFAR10、Tiny-ImageNet和TSRD数据集上,当投毒率分别仅为训练集规模的0.1%、0.025%和0.4%时,毒箭蛙相较于LC、HTBA、BadNets和Blend方法实现了更高的攻击成功率。此外,与最先进的攻击方法NARCISSUS相比,毒箭蛙在无需任何训练数据的情况下实现了相似的攻击成功率。最后,我们证明四种典型后门防御算法难以抵御毒箭蛙攻击。