Deep Neural Networks are powerful tools to understand complex patterns and making decisions. However, their black-box nature impedes a complete understanding of their inner workings. While online saliency-guided training methods try to highlight the prominent features in the model's output to alleviate this problem, it is still ambiguous if the visually explainable features align with robustness of the model against adversarial examples. In this paper, we investigate the saliency trained model's vulnerability to adversarial examples methods. Models are trained using an online saliency-guided training method and evaluated against popular algorithms of adversarial examples. We quantify the robustness and conclude that despite the well-explained visualizations in the model's output, the salient models suffer from the lower performance against adversarial examples attacks.
翻译:深度神经网络是理解复杂模式和做出决策的强大工具。然而,其黑箱特性阻碍了对其内部运作的完整理解。尽管在线显著性引导训练方法试图通过突出模型输出中的显著特征来缓解这一问题,但视觉上可解释的特征是否与模型对抗对抗样本的鲁棒性一致,仍不明确。本文研究了经显著性训练的模型对抗对抗样本方法的脆弱性。模型采用在线显著性引导训练方法进行训练,并针对流行的对抗样本算法进行评估。我们量化了鲁棒性并得出结论:尽管模型输出具有解释性良好的可视化特征,但显著性模型在对抗对抗样本攻击时性能较低。