The significant rise of security concerns in conventional centralized learning has promoted federated learning (FL) adoption in building intelligent applications without privacy breaches. In cybersecurity, the sensitive data along with the contextual information and high-quality labeling in each enterprise organization play an essential role in constructing high-performance machine learning (ML) models for detecting cyber threats. Nonetheless, the risks coming from poisoning internal adversaries against FL systems have raised discussions about designing robust anti-poisoning frameworks. Whereas defensive mechanisms in the past were based on outlier detection, recent approaches tend to be more concerned with latent space representation. In this paper, we investigate a novel robust aggregation method for FL, namely Fed-LSAE, which takes advantage of latent space representation via the penultimate layer and Autoencoder to exclude malicious clients from the training process. The experimental results on the CIC-ToN-IoT and N-BaIoT datasets confirm the feasibility of our defensive mechanism against cutting-edge poisoning attacks for developing a robust FL-based threat detector in the context of IoT. More specifically, the FL evaluation witnesses an upward trend of approximately 98% across all metrics when integrating with our Fed-LSAE defense.
翻译:传统集中式学习中的安全问题显著增加,推动了联邦学习在不泄露隐私的情况下构建智能应用的应用。在网络安全领域,各企业组织中的敏感数据及其上下文信息与高质量标注,对于构建高性能网络威胁检测机器学习模型至关重要。然而,内部投毒攻击者对联邦学习系统构成的威胁引发了关于设计鲁棒抗投毒框架的讨论。早期防御机制基于异常值检测,而近期方法更关注潜在空间表征。本文提出一种名为Fed-LSAE的新型鲁棒聚合方法,该方法利用倒数第二层和自编码器获取潜在空间表征,将恶意客户端排除在训练过程之外。在CIC-ToN-IoT和N-BaIoT数据集上的实验结果证实,该防御机制能够有效应对前沿投毒攻击,为开发物联网环境下基于联邦学习的鲁棒威胁检测器提供了可行性。具体而言,在联邦学习评估中,集成Fed-LSAE防御后各指标均呈现约98%的上升趋势。