Malicious communication behavior is the network communication behavior generated by malware (bot-net, spyware, etc.) after victim devices are infected. Experienced adversaries often hide malicious information in HTTP traffic to evade detection. However, related detection methods have inadequate generalization ability because they are usually based on artificial feature engineering and outmoded datasets. In this paper, we propose an HTTP-based Malicious Communication traffic Detection Model (HMCD-Model) based on generated adversarial flows and hierarchical traffic features. HMCD-Model consists of two parts. The first is a generation algorithm based on WGAN-GP to generate HTTP-based malicious communication traffic for data enhancement. The second is a hybrid neural network based on CNN and LSTM to extract hierarchical spatial-temporal features of traffic. In addition, we collect and publish a dataset, HMCT-2020, which consists of large-scale malicious and benign traffic during three years (2018-2020). Taking the data in HMCT-2020(18) as the training set and the data in other datasets as the test set, the experimental results show that the HMCD-Model can effectively detect unknown HTTP-based malicious communication traffic. It can reach F1 = 98.66% in the dataset HMCT-2020(19-20), F1 = 90.69% in the public dataset CIC-IDS-2017, and F1 = 83.66% in the real traffic, which is 20+% higher than other representative methods on average. This validates that HMCD-Model has the ability to discover unknown HTTP-based malicious communication behavior.
翻译:恶意通信行为是指受害设备被感染后,恶意软件(如僵尸网络、间谍软件等)产生的网络通信行为。老练的攻击者常将恶意信息隐藏于HTTP流量中以逃避检测。然而,现有检测方法由于通常基于人工特征工程和过时数据集,存在泛化能力不足的问题。本文提出一种基于生成对抗流和分层流量特征的HTTP恶意通信流量检测模型(HMCD-Model)。该模型由两部分组成:第一部分是基于WGAN-GP的生成算法,用于生成HTTP恶意通信流量以实现数据增强;第二部分是基于CNN和LSTM的混合神经网络,用于提取流量的分层时空特征。此外,我们收集并发布了HMCT-2020数据集,包含三年间(2018-2020年)的大规模恶意与良性流量。以HMCT-2020(18)数据为训练集、其他数据集为测试集的实验结果表明,HMCD-Model能有效检测未知HTTP恶意通信流量:在HMCT-2020(19-20)数据集上F1值达98.66%,在公开数据集CIC-IDS-2017上F1值达90.69%,在真实流量中F1值达83.66%,平均较其他代表性方法提升20%以上。这验证了HMCD-Model具备发现未知HTTP恶意通信行为的能力。