Machine learning (ML) has become increasingly popular in network intrusion detection. However, ML-based solutions always respond regardless of whether the input data reflects known patterns, a common issue across safety-critical applications. While several proposals exist for detecting Out-Of-Distribution (OOD) in other fields, it remains unclear whether these approaches can effectively identify new forms of intrusions for network security. New attacks, not necessarily affecting overall distributions, are not guaranteed to be clearly OOD as instead, images depicting new classes are in computer vision. In this work, we investigate whether existing OOD detectors from other fields allow the identification of unknown malicious traffic. We also explore whether more discriminative and semantically richer embedding spaces within models, such as those created with contrastive learning and multi-class tasks, benefit detection. Our investigation covers a set of six OOD techniques that employ different detection strategies. These techniques are applied to models trained in various ways and subsequently exposed to unknown malicious traffic from the same and different datasets (network environments). Our findings suggest that existing detectors can identify a consistent portion of new malicious traffic, and that improved embedding spaces enhance detection. We also demonstrate that simple combinations of certain detectors can identify almost 100% of malicious traffic in our tested scenarios.
翻译:机器学习在网络入侵检测中日益普及。然而,基于机器学习的解决方案无论输入数据是否反映已知模式都会做出响应,这是安全关键应用中的常见问题。尽管其他领域已有多种分布外检测方案,但尚不清楚这些方法能否有效识别网络安全中的新型入侵形式。与计算机视觉中描绘新类别的图像不同,新的攻击未必会影响整体数据分布,因此无法保证其清晰的分布外特征。本研究探讨了其他领域的现有分布外检测器是否能够识别未知恶意流量,并进一步探究模型中更具判别性和语义丰富的嵌入空间(例如通过对比学习和多分类任务构建的嵌入空间)是否有助于检测。我们研究了六种采用不同检测策略的分布外技术,将这些技术应用于经过多种方式训练的模型,随后让模型暴露于来自相同和不同数据集(网络环境)的未知恶意流量中。结果表明,现有检测器能够识别相当一部分新型恶意流量,且改进的嵌入空间能提升检测效果。我们还证明,某些检测器的简单组合可以在测试场景中识别几乎100%的恶意流量。