Intrusion detection systems (IDSs) play a critical role in protecting billions of IoT devices from malicious attacks. However, the IDSs for IoT devices face inherent challenges of IoT systems, including the heterogeneity of IoT data/devices, the high dimensionality of training data, and the imbalanced data. Moreover, the deployment of IDSs on IoT systems is challenging, and sometimes impossible, due to the limited resources such as memory/storage and computing capability of typical IoT devices. To tackle these challenges, this article proposes a novel deep neural network/architecture called Constrained Twin Variational Auto-Encoder (CTVAE) that can feed classifiers of IDSs with more separable/distinguishable and lower-dimensional representation data. Additionally, in comparison to the state-of-the-art neural networks used in IDSs, CTVAE requires less memory/storage and computing power, hence making it more suitable for IoT IDS systems. Extensive experiments with the 11 most popular IoT botnet datasets show that CTVAE can boost around 1% in terms of accuracy and Fscore in detection attack compared to the state-of-the-art machine learning and representation learning methods, whilst the running time for attack detection is lower than 2E-6 seconds and the model size is lower than 1 MB. We also further investigate various characteristics of CTVAE in the latent space and in the reconstruction representation to demonstrate its efficacy compared with current well-known methods.
翻译:入侵检测系统(IDS)在保护数十亿物联网设备免受恶意攻击中发挥着关键作用。然而,针对物联网设备的IDS面临物联网系统固有的挑战,包括物联网数据/设备的异构性、训练数据的高维度以及数据的不平衡性。此外,由于典型物联网设备在内存/存储和计算能力等资源受限,在物联网系统上部署IDS极具挑战性,有时甚至不可行。为应对这些挑战,本文提出了一种新颖的深度神经网络/架构——约束型孪生变分自编码器(CTVAE),该架构能够为IDS的分类器提供更具可分性/可辨识性且维度更低的表征数据。与当前IDS领域最先进的神经网络相比,CTVAE所需内存/存储和计算能力更低,因而更适用于物联网入侵检测系统。基于11个最流行的物联网僵尸网络数据集的大量实验表明,与最先进的机器学习和表征学习方法相比,CTVAE在攻击检测的准确率和F值上可提升约1%,同时攻击检测运行时间低于2E-6秒,模型大小小于1 MB。我们还进一步研究了CTVAE在潜在空间和重构表征中的多种特性,以证明其相较于当前主流方法的优越性。