As the smartphone market leader, Android has been a prominent target for malware attacks. The number of malicious applications (apps) identified for it has increased continually over the past decade, creating an immense challenge for all parties involved. For market holders and researchers, in particular, the large number of samples has made manual malware detection unfeasible, leading to an influx of research that investigate Machine Learning (ML) approaches to automate this process. However, while some of the proposed approaches achieve high performance, rapidly evolving Android malware has made them unable to maintain their accuracy over time. This has created a need in the community to conduct further research, and build more flexible ML pipelines. Doing so, however, is currently hindered by a lack of systematic overview of the existing literature, to learn from and improve upon the existing solutions. Existing survey papers often focus only on parts of the ML process (e.g., data collection or model deployment), while omitting other important stages, such as model evaluation and explanation. In this paper, we address this problem with a review of 42 highly-cited papers, spanning a decade of research (from 2011 to 2021). We introduce a novel procedural taxonomy of the published literature, covering how they have used ML algorithms, what features they have engineered, which dimensionality reduction techniques they have employed, what datasets they have employed for training, and what their evaluation and explanation strategies are. Drawing from this taxonomy, we also identify gaps in knowledge and provide ideas for improvement and future work.
翻译:作为智能手机市场的领导者,安卓系统已成为恶意软件攻击的重点目标。过去十年间,安卓平台上被识别的恶意应用程序数量持续增长,给所有相关方带来了巨大挑战。尤其对市场持有者和研究人员而言,海量样本使得人工恶意软件检测变得不可行,由此催生了大量研究探索利用机器学习方法自动化该过程。然而,尽管部分方法取得了优异性能,但快速演变的安卓恶意软件使其准确性难以长期维持。这促使学界需要开展进一步研究,构建更灵活的机器学习流水线。然而,当前缺乏对现有文献的系统性综述来学习并改进已有解决方案。现有综述论文往往仅关注机器学习流程的局部环节(如数据采集或模型部署),而忽略了模型评估与解释等重要阶段。本文通过系统梳理42篇高被引论文(涵盖2011至2021年十年的研究)解决了这一问题。我们提出了一种新型程序化分类法,覆盖了已发表文献中机器学习算法的应用方式、特征工程方法、降维技术选择、训练数据集使用情况以及评估与解释策略。基于该分类法,我们还识别了当前知识体系中的空白,并为后续改进与未来研究方向提供了思路。