The growing dependence on machine learning in real-world applications emphasizes the importance of understanding and ensuring its safety. Backdoor attacks pose a significant security risk due to their stealthy nature and potentially serious consequences. Such attacks involve embedding triggers within a learning model with the intention of causing malicious behavior when an active trigger is present while maintaining regular functionality without it. This paper evaluates the effectiveness of any backdoor attack incorporating a constant trigger, by establishing tight lower and upper boundaries for the performance of the compromised model on both clean and backdoor test data. The developed theory answers a series of fundamental but previously underexplored problems, including (1) what are the determining factors for a backdoor attack's success, (2) what is the direction of the most effective backdoor attack, and (3) when will a human-imperceptible trigger succeed. Our derived understanding applies to both discriminative and generative models. We also demonstrate the theory by conducting experiments using benchmark datasets and state-of-the-art backdoor attack scenarios.
翻译:机器学习在现实世界应用中的依赖性日益增强,凸显了理解并确保其安全性的重要性。后门攻击因其隐蔽性和潜在严重后果而构成重大安全威胁。此类攻击涉及在学习模型中植入触发器,旨在当触发器激活时引发恶意行为,而在无触发器时维持正常功能。本文通过建立受损模型在干净测试数据和后门测试数据上性能的严格上下界,评估了包含恒定触发器的任意后门攻击的有效性。所提出的理论解答了一系列基础但此前探究不足的问题,包括:(1)决定后门攻击成功的关键因素是什么,(2)最有效后门攻击的方向是什么,以及(3)何时人类不可察觉的触发器能够成功。我们推导出的理解同时适用于判别式模型和生成式模型。我们还通过使用基准数据集和最新后门攻击场景进行实验来验证该理论。