Almost 50 years after the invention of SQL, injection attacks are still top-tier vulnerabilities of today's ICT systems. Consequently, SQLi detection is still an active area of research, where the most recent works incorporate machine learning techniques into the proposed solutions. In this work, we highlight the shortcomings of the previous ML-based results focusing on four aspects: the evaluation methods, the optimization of the model parameters, the distribution of utilized datasets, and the feature selection. Since no single work explored all of these aspects satisfactorily, we fill this gap and provide an in-depth and comprehensive empirical analysis. Moreover, we cross-validate the trained models by using data from other distributions. This aspect of ML models (trained for SQLi detection) was never studied. Yet, the sensitivity of the model's performance to this is crucial for any real-life deployment. Finally, we validate our findings on a real-world industrial SQLi dataset.
翻译:自SQL发明近50年来,注入攻击仍是当今ICT系统中顶尖级别的安全漏洞。因此,SQL注入检测依然是一个活跃的研究领域,最新工作多将机器学习技术融入所提出的解决方案中。本研究聚焦四大方面指出现有基于机器学习的研究成果的不足:评估方法、模型参数优化、所用数据集的分布特性以及特征选择。由于尚无单项研究能充分涵盖所有方面,我们填补这一空白并开展深入全面的实证分析。此外,我们通过使用其他分布的数据对训练模型进行交叉验证。这一针对SQL注入检测的机器学习模型(训练阶段)的研究维度此前从未被探讨,然而模型性能对此的敏感性对于任何实际部署而言都至关重要。最终,我们利用真实的工业级SQL注入数据集验证了研究成果。