Building an Adversarial Malware Dataset by Family and Type: Generation, Evasion, and Poisoning Evaluation

We present a dataset of adversarial malware samples derived from the public RawMal-TF collection of real-world malware binaries. Using a suite of adversarial malware generators, we construct two sets of adversarial PE files: 44,347 family-labelled samples and 33,596 type-labelled samples, achieving evasion rates of 98.35 % and 92.20 % against the EMBER classifier, respectively. Each adversarial binary is accompanied by detailed metadata, including EMBER scores and VirusTotal classifications. We further demonstrate the susceptibility of malware classification pipelines to data poisoning attacks through a series of training experiments. Injecting fully mislabelled adversarial samples representing only 0.5 % of the training data in the family-labelled dataset increases the evasion rate against the re-trained classifier from 26.1 % to 92.8 %. The dataset is publicly released to facilitate future research on adversarial malware, poisoning attacks, and the robustness of machine-learning-based malware detection systems.

翻译：我们提出一个来源于公开RawMal-TF真实世界恶意软件二进制文件的对抗性恶意软件样本数据集。通过使用一套对抗性恶意软件生成器，我们构建了两组对抗性PE文件：44,347个带家族标签的样本和33,596个带类型标签的样本，分别对EMBER分类器实现了98.35%和92.20%的逃逸率。每个对抗性二进制文件均附有详细元数据，包括EMBER评分和VirusTotal分类结果。通过一系列训练实验，我们进一步证明了恶意软件分类流程对数据投毒攻击的敏感性。在带家族标签的数据集中，仅注入占训练数据0.5%的完全错误标记对抗性样本，即可使针对重新训练分类器的逃逸率从26.1%提升至92.8%。本数据集已公开发布，以促进未来在对抗性恶意软件、投毒攻击以及基于机器学习的恶意软件检测系统鲁棒性方面的研究。

相关内容

数据集

关注 88

数据集，又称为资料集、数据集合或资料集合，是一种由数据所组成的集合。
Data set（或dataset）是一个数据的集合，通常以表格形式出现。每一列代表一个特定变量。每一行都对应于某一成员的数据集的问题。它列出的价值观为每一个变量，如身高和体重的一个物体或价值的随机数。每个数值被称为数据资料。对应于行数，该数据集的数据可能包括一个或多个成员。

《基于动态图神经网络的恶意软件检测》

专知会员服务

16+阅读 · 1月28日

《分布式机器人群体聚类：增加抵御恶意伪装攻击的能力》

专知会员服务

29+阅读 · 2023年11月6日

《基于对手网络基础设施发掘来实现自动威胁建模》2023最新79页论文

专知会员服务

33+阅读 · 2023年5月14日

《H4rm0ny：用于规避恶意软件生成和检测的多智能体学习的竞争性两人零和马尔可夫博弈》2022最新12页论文，加拿大国防研究与发展部

专知会员服务

27+阅读 · 2022年10月26日