This work presents an evaluation of six prominent commercial endpoint malware detectors, a network malware detector, and a file-conviction algorithm from a cyber technology vendor. The evaluation was administered as the first of the Artificial Intelligence Applications to Autonomous Cybersecurity (AI ATAC) prize challenges, funded by / completed in service of the US Navy. The experiment employed 100K files (50/50% benign/malicious) with a stratified distribution of file types, including ~1K zero-day program executables (increasing experiment size two orders of magnitude over previous work). We present an evaluation process of delivering a file to a fresh virtual machine donning the detection technology, waiting 90s to allow static detection, then executing the file and waiting another period for dynamic detection; this allows greater fidelity in the observational data than previous experiments, in particular, resource and time-to-detection statistics. To execute all 800K trials (100K files $\times$ 8 tools), a software framework is designed to choreographed the experiment into a completely automated, time-synced, and reproducible workflow with substantial parallelization. A cost-benefit model was configured to integrate the tools' recall, precision, time to detection, and resource requirements into a single comparable quantity by simulating costs of use. This provides a ranking methodology for cyber competitions and a lens through which to reason about the varied statistical viewpoints of the results. These statistical and cost-model results provide insights on state of commercial malware detection.
翻译:本文评估了来自一家网络安全技术供应商的六款主流商业端点恶意软件检测器、一款网络恶意软件检测器以及一个文件定罪算法。该评估作为人工智能自主网络安全(AI ATAC)挑战赛的首项任务开展,由美国海军资助并为其服务。实验采用了10万个文件(良性/恶意文件各占50%),按文件类型进行分层分布,其中包括约1000个零日程序可执行文件(实验规模较之前研究提升了两个数量级)。我们提出了一种评估流程:将文件传输至部署了检测技术的新虚拟机,等待90秒进行静态检测,随后执行文件并再等待一段时间进行动态检测;这使得观测数据(尤其是资源消耗和检测时间统计数据)比以往实验具有更高保真度。为完成全部80万次试验(10万个文件×8种工具),我们设计了一个软件框架,将实验编排为完全自动化、时间同步且可复现的工作流程,并实现了大幅并行化处理。通过模拟使用成本,我们配置了一个成本效益模型,将工具的召回率、精确率、检测时间和资源需求整合为单一可比数值。这为网络竞赛提供了排名方法论,也为从不同统计视角解读实验结果提供了分析框架。这些统计与成本模型结果揭示了商业恶意软件检测技术的发展现状。