The origin of adversarial examples is still inexplicable in research fields, and it arouses arguments from various viewpoints, albeit comprehensive investigations. In this paper, we propose a way of delving into the unexpected vulnerability in adversarially trained networks from a causal perspective, namely adversarial instrumental variable (IV) regression. By deploying it, we estimate the causal relation of adversarial prediction under an unbiased environment dissociated from unknown confounders. Our approach aims to demystify inherent causal features on adversarial examples by leveraging a zero-sum optimization game between a casual feature estimator (i.e., hypothesis model) and worst-case counterfactuals (i.e., test function) disturbing to find causal features. Through extensive analyses, we demonstrate that the estimated causal features are highly related to the correct prediction for adversarial robustness, and the counterfactuals exhibit extreme features significantly deviating from the correct prediction. In addition, we present how to effectively inoculate CAusal FEatures (CAFE) into defense networks for improving adversarial robustness.
翻译:对抗样本的成因在学术研究中仍难以解释,尽管已有全面探讨,但各派观点仍存争议。本文从因果视角提出一种探究对抗训练网络中意外脆弱性的方法,即对抗工具变量回归。通过部署该方法,我们能够在与未知混杂因素无关的无偏环境下估计对抗预测的因果关系。我们的方法旨在通过构建因果特征估计器(即假设模型)与干扰因果特征发现的极端反事实(即测试函数)之间的零和博弈,揭示对抗样本中固有的因果特征。通过大量分析,我们证明了所估计的因果特征与对抗鲁棒性的正确预测高度相关,而反事实则展现出显著偏离正确预测的极端特征。此外,我们提出了如何有效将因果特征注入防御网络以提升对抗鲁棒性。