HTTP-based Trojan is extremely threatening, and it is difficult to be effectively detected because of its concealment and confusion. Previous detection methods usually are with poor generalization ability due to outdated datasets and reliance on manual feature extraction, which makes these methods always perform well under their private dataset, but poorly or even fail to work in real network environment. In this paper, we propose an HTTP-based Trojan detection model via the Hierarchical Spatio-Temporal Features of traffics (HSTF-Model) based on the formalized description of traffic spatio-temporal behavior from both packet level and flow level. In this model, we employ Convolutional Neural Network (CNN) to extract spatial information and Long Short-Term Memory (LSTM) to extract temporal information. In addition, we present a dataset consisting of Benign and Trojan HTTP Traffic (BTHT-2018). Experimental results show that our model can guarantee high accuracy (the F1 of 98.62%-99.81% and the FPR of 0.34%-0.02% in BTHT-2018). More importantly, our model has a huge advantage over other related methods in generalization ability. HSTF-Model trained with BTHT-2018 can reach the F1 of 93.51% on the public dataset ISCX-2012, which is 20+% better than the best of related machine learning methods.
翻译:HTTP木马极具威胁性,因其隐蔽性和混淆特性难以被有效检测。现有检测方法通常因数据集陈旧且依赖人工特征提取而泛化能力差,导致在私有数据集上表现良好,但在真实网络环境中效果不佳甚至完全失效。本文基于数据包级和流级流量的时空行为形式化描述,提出一种基于流量层次时空特征的HTTP木马检测模型(HSTF-Model)。该模型采用卷积神经网络(CNN)提取空间信息,并利用长短期记忆网络(LSTM)提取时间信息。此外,我们构建了一个包含良性HTTP流量和木马HTTP流量的数据集(BTHT-2018)。实验结果表明,该模型能保证高准确率(在BTHT-2018上F1值达98.62%-99.81%,假阳性率(FPR)为0.34%-0.02%)。更重要的是,本模型在泛化能力上显著优于其他相关方法:使用BTHT-2018训练的HSTF-Model在公开数据集ISCX-2012上F1值可达93.51%,比最优的机器学习方法高出20%以上。