Adversarial machine learning (AML) studies the adversarial phenomenon of machine learning, which may make inconsistent or unexpected predictions with humans. Some paradigms have been recently developed to explore this adversarial phenomenon occurring at different stages of a machine learning system, such as training-time adversarial attack (i.e., backdoor attack), deployment-time adversarial attack (i.e., weight attack), and inference-time adversarial attack (i.e., adversarial example). However, although these paradigms share a common goal, their developments are almost independent, and there is still no big picture of AML. In this work, we aim to provide a unified perspective to the AML community to systematically review the overall progress of this field. We firstly provide a general definition about AML, and then propose a unified mathematical framework to covering existing attack paradigms. According to the proposed unified framework, we can not only clearly figure out the connections and differences among these paradigms, but also systematically categorize and review existing works in each paradigm.
翻译:对抗性机器学习(AML)研究机器学习中可能产生与人类不一致或不可预测预测结果的对抗现象。近期,一些研究范式被提出以探索机器学习系统不同阶段出现的此类对抗现象,包括训练时对抗攻击(即后门攻击)、部署时对抗攻击(即权重攻击)以及推断时对抗攻击(即对抗样本)。然而,尽管这些范式具有共同目标,其发展几乎相互独立,且目前仍缺乏对AML的全局性认知。本文旨在为AML社区提供统一视角,系统综述该领域的整体进展。我们首先给出AML的一般性定义,继而提出统一的数学框架以覆盖现有攻击范式。基于该统一框架,我们不仅能够清晰厘清这些范式间的关联与差异,还能系统地对各范式下的现有工作进行分类与综述。