We present a new algorithm to train a robust malware detector. Modern malware detectors rely on machine learning algorithms. Now, the adversarial objective is to devise alterations to the malware code to decrease the chance of being detected whilst preserving the functionality and realism of the malware. Adversarial learning is effective in improving robustness but generating functional and realistic adversarial malware samples is non-trivial. Because: i) in contrast to tasks capable of using gradient-based feedback, adversarial learning in a domain without a differentiable mapping function from the problem space (malware code inputs) to the feature space is hard; and ii) it is difficult to ensure the adversarial malware is realistic and functional. This presents a challenge for developing scalable adversarial machine learning algorithms for large datasets at a production or commercial scale to realize robust malware detectors. We propose an alternative; perform adversarial learning in the feature space in contrast to the problem space. We prove the projection of perturbed, yet valid malware, in the problem space into feature space will always be a subset of adversarials generated in the feature space. Hence, by generating a robust network against feature-space adversarial examples, we inherently achieve robustness against problem-space adversarial examples. We formulate a Bayesian adversarial learning objective that captures the distribution of models for improved robustness. We prove that our learning method bounds the difference between the adversarial risk and empirical risk explaining the improved robustness. We show that adversarially trained BNNs achieve state-of-the-art robustness. Notably, adversarially trained BNNs are robust against stronger attacks with larger attack budgets by a margin of up to 15% on a recent production-scale malware dataset of more than 20 million samples.
翻译:我们提出了一种新的算法来训练鲁棒的恶意软件检测器。现代恶意软件检测器依赖于机器学习算法。当前,对抗目标是通过对恶意软件代码进行修改,降低其被检测概率,同时保留恶意软件的功能性和真实性。对抗学习在提升鲁棒性方面是有效的,但生成功能性和真实性的对抗性恶意软件样本并非易事。原因在于:i) 与能够利用基于梯度的反馈任务不同,在缺乏从问题空间(恶意软件代码输入)到特征空间的可微分映射函数的领域中进行对抗学习较为困难;ii) 难以确保对抗性恶意软件的真实性和功能性。这为在工业生产或商业规模的大规模数据集上开发可扩展的对抗机器学习算法以实现鲁棒恶意软件检测器带来了挑战。我们提出了一种替代方案:在特征空间而非问题空间中进行对抗学习。我们证明了问题空间中经扰动但仍有效的恶意软件投影到特征空间后,始终是特征空间中所生成对抗样本的子集。因此,通过训练针对特征空间对抗样本的鲁棒网络,我们固有地实现了针对问题空间对抗样本的鲁棒性。我们构建了一个贝叶斯对抗学习目标,以捕获模型分布从而提升鲁棒性。我们证明,我们的学习方法界定了对抗风险与经验风险之间的差异,从而解释了鲁棒性的提升。我们展示了经对抗训练的贝叶斯神经网络(BNN)达到了最先进的鲁棒性。值得注意的是,在最近一个包含超过2000万个样本的生产规模恶意软件数据集上,经对抗训练的BNN能够以高达15%的余量抵御具有更大攻击预算的更强攻击。