Many organisations manage service quality and monitor a large set devices and servers where each entity is associated with telemetry or physical sensor data series. Recently, various methods have been proposed to detect behavioural anomalies, however existing approaches focus on multivariate time series and ignore communication between entities. Moreover, we aim to support end-users in not only in locating entities and sensors causing an anomaly at a certain period, but also explain this decision. We propose a scalable approach to detect anomalies using a two-step approach. First, we recover relations between entities in the network, since relations are often dynamic in nature and caused by an unknown underlying process. Next, we report anomalies based on an embedding of sequential patterns. Pattern mining is efficient and supports interpretation, i.e. patterns represent frequent occurring behaviour in time series. We extend pattern mining to filter sequential patterns based on frequency, temporal constraints and minimum description length. We collect and release two public datasets for international broadcasting and X from an Internet company. \textit{BAD} achieves an overall F1-Score of 0.78 on 9 benchmark datasets, significantly outperforming the best baseline by 3\%. Additionally, \textit{BAD} is also an order-of-magnitude faster than state-of-the-art anomaly detection methods.
翻译:许多组织管理服务质量并监控大量设备和服务器,其中每个实体关联遥测或物理传感器数据序列。近年来,已有多种方法被提出用于检测行为异常,但现有方法主要关注多变量时间序列,忽略了实体间的通信关系。此外,我们不仅希望帮助终端用户定位特定时间段内引发异常的实体和传感器,还希望解释这一决策。我们提出一种可扩展的两步式异常检测方法。首先,由于网络实体间的关系常具有动态特性且由未知底层过程驱动,我们恢复实体间的关联关系。其次,我们基于序列模式的嵌入表示报告异常。模式挖掘具有高效性且支持可解释性,即模式代表时间序列中频繁出现的行为。我们将模式挖掘扩展为基于频率、时间约束和最小描述长度过滤序列模式的方法。我们收集并发布了两个面向国际广播和某互联网公司X的公开数据集。实验表明,\textit{BAD}方法在9个基准数据集上取得了0.78的整体F1分数,显著优于最佳基线方法3\%。此外,\textit{BAD}在速度上比最先进的异常检测方法快一个数量级。