Typical event datasets such as those used in network intrusion detection comprise hundreds of thousands, sometimes millions, of discrete packet events. These datasets tend to be high dimensional, stateful, and time-series in nature, holding complex local and temporal feature associations. Packet data can be abstracted into lower dimensional summary data, such as packet flow records, where some of the temporal complexities of packet data can be mitigated, and smaller well-engineered feature subsets can be created. This data can be invaluable as training data for machine learning and cyber threat detection techniques. Data can be collected in real-time, or from historical packet trace archives. In this paper we focus on how flow records and summary metadata can be extracted from packet data with high accuracy and robustness. We identify limitations in current methods, how they may impact datasets, and how these flaws may impact learning models. Finally, we propose methods to improve the state of the art and introduce proof of concept tools to support this work.
翻译:典型的事件数据集(如网络入侵检测中使用的数据集)包含数十万乃至数百万个离散数据包事件。这些数据集本质上具有高维度、有状态和时间序列特性,包含复杂的局部和时序特征关联。数据包数据可抽象为低维度的概要数据(如数据包流记录),从而缓解数据包数据的部分时序复杂性,并生成经过精心设计的小规模特征子集。这类数据可作为机器学习和网络威胁检测技术的重要训练数据源,其既可通过实时采集获取,也可从历史数据包跟踪档案中提取。本文重点研究如何从数据包数据中高精度、高鲁棒性地提取流记录和概要元数据。我们指出现有方法的局限性及其对数据集可能产生的影响,并分析这些缺陷如何影响学习模型。最后,我们提出改进现有技术的方案,并引入支持本项工作的概念验证工具。