The remarkable success of the use of machine learning-based solutions for network security problems has been impeded by the developed ML models' inability to maintain efficacy when used in different network environments exhibiting different network behaviors. This issue is commonly referred to as the generalizability problem of ML models. The community has recognized the critical role that training datasets play in this context and has developed various techniques to improve dataset curation to overcome this problem. Unfortunately, these methods are generally ill-suited or even counterproductive in the network security domain, where they often result in unrealistic or poor-quality datasets. To address this issue, we propose an augmented ML pipeline that leverages explainable ML tools to guide the network data collection in an iterative fashion. To ensure the data's realism and quality, we require that the new datasets should be endogenously collected in this iterative process, thus advocating for a gradual removal of data-related problems to improve model generalizability. To realize this capability, we develop a data-collection platform, netUnicorn, that takes inspiration from the classic "hourglass" model and is implemented as its "thin waist" to simplify data collection for different learning problems from diverse network environments. The proposed system decouples data-collection intents from the deployment mechanisms and disaggregates these high-level intents into smaller reusable, self-contained tasks. We demonstrate how netUnicorn simplifies collecting data for different learning problems from multiple network environments and how the proposed iterative data collection improves a model's generalizability.
翻译:基于机器学习的解决方案在网络安全问题中取得了显著成功,但其开发出的模型在不同网络环境中因网络行为差异而无法保持有效性,这一问题被普遍称为机器学习模型的泛化性问题。学界已认识到训练数据集在此背景下的关键作用,并开发了多种改进数据集策管的技术以解决该问题。遗憾的是,这些方法在网络安全领域中通常不适用甚至适得其反,往往导致数据集不真实或质量低下。为应对这一挑战,我们提出一种增强型机器学习流水线,利用可解释机器学习工具迭代式引导网络数据采集。为确保数据的真实性与质量,我们要求在此迭代过程中内生性地采集新数据集,从而倡导逐步消除数据相关问题以提升模型泛化能力。为实现这一能力,我们开发了数据采集平台netUnicorn,该平台借鉴经典"沙漏"模型理念,作为其"细腰"组件实现,以简化从不同网络环境中采集面向不同学习问题的数据。该平台将数据采集意图与部署机制解耦,并将这些高层意图分解为可复用、自包含的微型任务。我们通过实验验证了netUnicorn如何简化从多个网络环境采集面向不同学习问题的数据,以及所提出的迭代式数据采集如何提升模型泛化性。