In this work, we conduct a comprehensive study on the robustness of domain generation algorithm (DGA) classifiers. We implement 32 white-box attacks, 19 of which are very effective and induce a false-negative rate (FNR) of $\approx$ 100\% on unhardened classifiers. To defend the classifiers, we evaluate different hardening approaches and propose a novel training scheme that leverages adversarial latent space vectors and discretized adversarial domains to significantly improve robustness. In our study, we highlight a pitfall to avoid when hardening classifiers and uncover training biases that can be easily exploited by attackers to bypass detection, but which can be mitigated by adversarial training (AT). In our study, we do not observe any trade-off between robustness and performance, on the contrary, hardening improves a classifier's detection performance for known and unknown DGAs. We implement all attacks and defenses discussed in this paper as a standalone library, which we make publicly available to facilitate hardening of DGA classifiers: https://gitlab.com/rwth-itsec/robust-dga-detection
翻译:在本文中,我们对域生成算法(DGA)分类器的鲁棒性进行了全面研究。我们实现了32种白盒攻击,其中19种攻击非常有效,能在未加固的分类器上诱导出约100%的假阴性率(FNR)。为防御这些攻击,我们评估了不同的加固方法,并提出了一种新颖的训练方案,利用对抗性潜在空间向量和离散化对抗域显著提升鲁棒性。在研究中,我们揭示了加固分类器时应避免的一个陷阱,并发现了可被攻击者轻易利用以绕过检测的训练偏差,而对抗训练(AT)可缓解这些偏差。我们未观察到鲁棒性与性能之间存在任何权衡,相反,加固能提升分类器对已知和未知DGA的检测性能。我们将本文讨论的所有攻击与防御实现为一个独立库,并公开发布以促进DGA分类器的加固:https://gitlab.com/rwth-itsec/robust-dga-detection