The growing dependence on machine learning in real-world applications emphasizes the importance of understanding and ensuring its safety. Backdoor attacks pose a significant security risk due to their stealthy nature and potentially serious consequences. Such attacks involve embedding triggers within a learning model with the intention of causing malicious behavior when an active trigger is present while maintaining regular functionality without it. This paper evaluates the effectiveness of any backdoor attack incorporating a constant trigger, by establishing tight lower and upper boundaries for the performance of the compromised model on both clean and backdoor test data. The developed theory answers a series of fundamental but previously underexplored problems, including (1) what are the determining factors for a backdoor attack's success, (2) what is the direction of the most effective backdoor attack, and (3) when will a human-imperceptible trigger succeed. Our derived understanding applies to both discriminative and generative models. We also demonstrate the theory by conducting experiments using benchmark datasets and state-of-the-art backdoor attack scenarios.
翻译:现实世界应用对机器学习的日益依赖,凸显了理解并确保其安全性的重要性。后门攻击因其隐蔽性与潜在严重后果,构成了重大安全风险。此类攻击通过在模型中嵌入触发器,使其在触发条件激活时产生恶意行为,而在无触发时维持正常功能。本文建立了一个下界与上界均严格的理论框架,用于评估携带恒定触发器的后门攻击效果,该框架适用于干净数据与后门测试数据上的被攻击模型性能。所提出的理论回答了一系列基础但此前未被充分探索的问题,包括:(1) 决定后门攻击成功的关键因素是什么;(2) 最有效的后门攻击方向是什么;(3) 人眼不可见的触发器在何种条件下能够成功。我们推导的结论同时适用于判别模型与生成模型。通过使用基准数据集与最先进的后门攻击场景进行实验,进一步验证了该理论。