Adversarial Machine Learning (AML) is emerging as a major field aimed at protecting machine learning (ML) systems against security threats: in certain scenarios there may be adversaries that actively manipulate input data to fool learning systems. This creates a new class of security vulnerabilities that ML systems may face, and a new desirable property called adversarial robustness essential to trust operations based on ML outputs. Most work in AML is built upon a game-theoretic modelling of the conflict between a learning system and an adversary, ready to manipulate input data. This assumes that each agent knows their opponent's interests and uncertainty judgments, facilitating inferences based on Nash equilibria. However, such common knowledge assumption is not realistic in the security scenarios typical of AML. After reviewing such game-theoretic approaches, we discuss the benefits that Bayesian perspectives provide when defending ML-based systems. We demonstrate how the Bayesian approach allows us to explicitly model our uncertainty about the opponent's beliefs and interests, relaxing unrealistic assumptions, and providing more robust inferences. We illustrate this approach in supervised learning settings, and identify relevant future research problems.
翻译:对抗性机器学习(AML)正逐渐成为保护机器学习系统免受安全威胁的重要研究领域:在某些场景中,可能存在主动操纵输入数据以欺骗学习系统的对抗者。这引发了机器学习系统可能面临的一类新型安全漏洞,以及一项名为对抗鲁棒性的理想特性,该特性对于基于机器学习输出的可信操作至关重要。大多数AML研究建立在学习系统与对抗者之间冲突的博弈论建模基础上,其中对抗者准备操纵输入数据。这种建模假设每个参与者都知晓对手的目标和不确定性判断,从而便于基于纳什均衡进行推理。然而,这种共有知识假设在AML典型的安全场景中并不现实。在综述这类博弈论方法后,我们探讨了贝叶斯视角在防御基于机器学习系统时带来的优势。我们展示了贝叶斯方法如何显式建模我们对于对手信念和目标的不确定性,放宽不切实际的假设,并提供更稳健的推理结果。我们以监督学习场景为例阐述该方法,并识别出相关的未来研究方向。