Distribution shift, a change in the statistical properties of data over time, poses a critical challenge for deep learning anomaly detection systems. Existing anomaly detection systems often struggle to adapt to these shifts. Specifically, systems based on supervised learning require costly manual labeling, while those based on unsupervised learning rely on clean data, which is difficult to obtain, for shift adaptation. Both of these requirements are challenging to meet in practice. In this paper, we introduce NetSight, a framework for supervised anomaly detection in network data that continually detects and adapts to distribution shifts in an online manner. NetSight eliminates manual intervention through a novel pseudo-labeling technique and uses a knowledge distillation-based adaptation strategy to prevent catastrophic forgetting. Evaluated on three long-term network datasets, NetSight demonstrates superior adaptation performance compared to state-of-the-art methods that rely on manual labeling, achieving F1-score improvements of up to 11.72%. This proves its robustness and effectiveness in dynamic networks that experience distribution shifts over time.
翻译:分布漂移(数据统计属性随时间的变化)对基于深度学习的异常检测系统构成了严峻挑战。现有异常检测系统往往难以适应这类漂移。具体而言,基于监督学习的系统需要昂贵的人工标注,而基于无监督学习的系统则依赖难以获取的干净数据进行漂移适应。这两种要求在现实中都难以满足。本文提出NetSight框架——一种网络数据上的监督异常检测方法,能够持续在线检测并适应分布漂移。NetSight通过新型伪标签技术消除人工干预,并采用知识蒸馏的自适应策略防止灾难性遗忘。在三个长期网络数据集上的评估表明,相较于依赖人工标注的最先进方法,NetSight展现出更优的自适应性能,F1分数提升高达11.72%。这证明了其在经历分布漂移的动态网络中的鲁棒性和有效性。