Anomaly detection, where data instances are discovered containing feature patterns different from the majority, plays a fundamental role in various applications. However, it is challenging for existing methods to handle the scenarios where the instances are systems whose characteristics are not readily observed as data. Appropriate interactions are needed to interact with the systems and identify those with abnormal responses. Detecting system-wise anomalies is a challenging task due to several reasons including: how to formally define the system-wise anomaly detection problem; how to find the effective activation signal for interacting with systems to progressively collect the data and learn the detector; how to guarantee stable training in such a non-stationary scenario with real-time interactions? To address the challenges, we propose InterSAD (Interactive System-wise Anomaly Detection). Specifically, first, we adopt Markov decision process to model the interactive systems, and define anomalous systems as anomalous transition and anomalous reward systems. Then, we develop an end-to-end approach which includes an encoder-decoder module that learns system embeddings, and a policy network to generate effective activation for separating embeddings of normal and anomaly systems. Finally, we design a training method to stabilize the learning process, which includes a replay buffer to store historical interaction data and allow them to be re-sampled. Experiments on two benchmark environments, including identifying the anomalous robotic systems and detecting user data poisoning in recommendation models, demonstrate the superiority of InterSAD compared with state-of-the-art baselines methods.
翻译:异常检测旨在发现那些特征模式与大多数数据不同的数据实例,在各类应用中发挥着基础性作用。然而,现有方法难以应对实例为系统且其特性无法直接观测为数据的场景。需要设计合适的交互方式与系统进行交互,并识别出具有异常响应的系统。检测系统级异常是一项具有挑战性的任务,原因包括:如何正式定义系统级异常检测问题;如何寻找有效的激活信号与系统交互,逐步收集数据并学习检测器;如何在实时交互的非平稳场景中保证稳定的训练过程?为解决这些挑战,我们提出了InterSAD(交互式系统级异常检测方法)。具体而言,首先采用马尔可夫决策过程对交互系统进行建模,并将异常系统定义为异常转移系统和异常奖励系统。其次,开发了一种端到端方法,包含学习系统嵌入的编码器-解码器模块,以及生成有效激活信号以分离正常与异常系统嵌入的策略网络。最后,设计了一种训练方法来稳定学习过程,包括使用经验回放缓冲区存储历史交互数据并允许其重采样。在两个基准环境(包括识别异常机器人系统以及检测推荐模型中的用户数据投毒)上的实验表明,InterSAD相比最先进的基线方法具有优越性。