Advance Persistent Threats (APTs), adopted by most delicate attackers, are becoming increasing common and pose great threat to various enterprises and institutions. Data provenance analysis on provenance graphs has emerged as a common approach in APT detection. However, previous works have exhibited several shortcomings: (1) requiring attack-containing data and a priori knowledge of APTs, (2) failing in extracting the rich contextual information buried within provenance graphs and (3) becoming impracticable due to their prohibitive computation overhead and memory consumption. In this paper, we introduce MAGIC, a novel and flexible self-supervised APT detection approach capable of performing multi-granularity detection under different level of supervision. MAGIC leverages masked graph representation learning to model benign system entities and behaviors, performing efficient deep feature extraction and structure abstraction on provenance graphs. By ferreting out anomalous system behaviors via outlier detection methods, MAGIC is able to perform both system entity level and batched log level APT detection. MAGIC is specially designed to handle concept drift with a model adaption mechanism and successfully applies to universal conditions and detection scenarios. We evaluate MAGIC on three widely-used datasets, including both real-world and simulated attacks. Evaluation results indicate that MAGIC achieves promising detection results in all scenarios and shows enormous advantage over state-of-the-art APT detection approaches in performance overhead.
翻译:先进持续性威胁(APT)作为顶尖攻击者最常采用的手段,正日益普遍,对各类企业和机构构成重大威胁。基于溯源图的数据来源分析已成为APT检测的常见方法。然而,现有研究存在以下不足:(1)需要包含攻击数据及APT先验知识,(2)无法有效提取溯源图中蕴含的丰富上下文信息,(3)因计算开销和内存消耗过高而难以落地实施。本文提出MAGIC——一种新颖且灵活的自监督APT检测方法,能够在不同监督级别下实现多粒度检测。MAGIC利用掩码图表示学习对良性系统实体及行为进行建模,对溯源图实现高效的深层特征提取与结构抽象。通过异常检测方法识别系统异常行为,MAGIC能够在系统实体级和批量日志级分别实现APT检测。该方法特别设计了模型自适应机制以处理概念漂移问题,并成功适用于通用场景及多种检测环境。我们在三个广泛使用的数据集上(涵盖真实攻击与模拟攻击)进行了评估。结果表明,MAGIC在所有场景下均展现出优异的检测效果,并在性能开销上较现有最先进APT检测方法具有显著优势。