Deep Learning (DL) is increasingly used in safety-critical applications, raising concerns about its reliability. DL suffers from a well-known problem of lacking robustness, especially when faced with adversarial perturbations known as Adversarial Examples (AEs). Despite recent efforts to detect AEs using advanced attack and testing methods, these approaches often overlook the input distribution and perceptual quality of the perturbations. As a result, the detected AEs may not be relevant in practical applications or may appear unrealistic to human observers. This can waste testing resources on rare AEs that seldom occur during real-world use, limiting improvements in DL model dependability. In this paper, we propose a new robustness testing approach for detecting AEs that considers both the feature level distribution and the pixel level distribution, capturing the perceptual quality of adversarial perturbations. The two considerations are encoded by a novel hierarchical mechanism. First, we select test seeds based on the density of feature level distribution and the vulnerability of adversarial robustness. The vulnerability of test seeds are indicated by the auxiliary information, that are highly correlated with local robustness. Given a test seed, we then develop a novel genetic algorithm based local test case generation method, in which two fitness functions work alternatively to control the perceptual quality of detected AEs. Finally, extensive experiments confirm that our holistic approach considering hierarchical distributions is superior to the state-of-the-arts that either disregard any input distribution or only consider a single (non-hierarchical) distribution, in terms of not only detecting imperceptible AEs but also improving the overall robustness of the DL model under testing.
翻译:深度学习(DL)正越来越多地应用于安全关键领域,这引发了对其可靠性的担忧。DL存在一个众所周知的鲁棒性缺失问题,尤其是在面对被称为对抗样本(AEs)的对抗性扰动时。尽管近年来人们尝试使用先进的攻击与测试方法来检测对抗样本,但这些方法往往忽略了输入分布及扰动的感知质量。因此,检测到的对抗样本在实际应用中可能不相关,或对人类观察者而言显得不真实。这会浪费测试资源在现实使用中极少出现的罕见对抗样本上,限制了DL模型可靠性的提升。本文提出一种新的鲁棒性测试方法用于检测对抗样本,该方法同时考虑了特征级分布和像素级分布,从而捕捉对抗扰动的感知质量。这两种考量通过一种新颖的层次化机制进行编码。首先,我们基于特征级分布的密度和对抗鲁棒性的脆弱性来选择测试种子。测试种子的脆弱性由与局部鲁棒性高度相关的辅助信息指示。给定测试种子后,我们随后开发了一种基于遗传算法的局部测试用例生成方法,其中两个适应度函数交替工作以控制检测到的对抗样本的感知质量。最后,大量实验证实,我们这种考虑层次化分布的整体方法在检测不可感知的对抗样本以及提升被测DL模型的整体鲁棒性方面,均优于忽视输入分布或仅考虑单一(非层次化)分布的最新方法。