Federated Learning (FL) enables collaborative deep learning training across multiple participants without exposing sensitive personal data. However, the distributed nature of FL and the unvetted participants' data makes it vulnerable to backdoor attacks. In these attacks, adversaries inject malicious functionality into the centralized model during training, leading to intentional misclassifications for specific adversary-chosen inputs. While previous research has demonstrated successful injections of persistent backdoors in FL, the persistence also poses a challenge, as their existence in the centralized model can prompt the central aggregation server to take preventive measures to penalize the adversaries. Therefore, this paper proposes a methodology that enables adversaries to effectively remove backdoors from the centralized model upon achieving their objectives or upon suspicion of possible detection. The proposed approach extends the concept of machine unlearning and presents strategies to preserve the performance of the centralized model and simultaneously prevent over-unlearning of information unrelated to backdoor patterns, making the adversaries stealthy while removing backdoors. To the best of our knowledge, this is the first work that explores machine unlearning in FL to remove backdoors to the benefit of adversaries. Exhaustive evaluation considering image classification scenarios demonstrates the efficacy of the proposed method in efficient backdoor removal from the centralized model, injected by state-of-the-art attacks across multiple configurations.
翻译:联邦学习(FL)支持多个参与方在不暴露敏感个人数据的情况下进行协作式深度学习训练。然而,FL的分布式特性以及未经验证的参与方数据使其易受后门攻击。在这些攻击中,攻击者在训练期间向中央模型中注入恶意功能,导致对特定攻击者选定输入的故意误分类。先前的研究已成功在FL中注入持久后门,但持久性也带来了挑战——因为后门在中央模型中的存在可能促使中央聚合服务器采取预防措施来惩罚攻击者。因此,本文提出一种方法,使攻击者能够在达成目标或怀疑可能被检测到时,有效地从中央模型中移除后门。所提方法扩展了机器遗忘的概念,并提出了保留中央模型性能同时防止后门模式无关信息被过度遗忘的策略,使攻击者在移除后门时保持隐蔽。据我们所知,这是首个探索在FL中利用机器遗忘移除后门以利于攻击者的研究。针对图像分类场景的全面评估表明,所提方法能在多种配置下有效移除由最先进攻击注入到中央模型中的后门。