Malware detection remains largely reactive: machine learning models trained on known samples degrade as threats evolve. Understanding evolutionary relationships among malware families can inform proactive defense, but traditional reverse engineering can take months to years to uncover such lineage relationships. We propose MalTree, a framework that applies bioinformatics inspired phylogenetic techniques (UPGMA and Neighbor-Joining) at scale to model malware evolution automatically using structural, behavioral, and image-based features. We introduce temporal validation using VirusTotal timestamps to assess whether inferred trees reflect actual evolutionary order. MalTree achieves 87% temporal consistency, indicating that inferred evolutionary relationships closely align with real-world emergence timelines. Our analysis shows that some families mutate over 10 times faster than others, suggesting that detection strategies should be tailored to family-specific evolutionary tempos. Case studies, including the Mirai botnet, confirm that inferred relationships from our phylogenetic tree align with documented threat intelligence. Our framework provides a foundation for shifting malware analysis from sample-by-sample classification toward lineage-aware evolutionary modeling.
翻译:恶意软件检测在很大程度上仍是被动的:基于已知样本训练的机器学习模型会随着威胁的演变而性能下降。理解恶意软件家族间的演化关系可为主动防御提供依据,但传统逆向工程可能需要数月乃至数年才能揭示此类谱系关系。我们提出MalTree框架,该框架受生物信息学启发,采用系统发育技术(UPGMA和邻接法)进行大规模自动化建模,利用结构、行为和基于图像的特征来分析恶意软件演化。我们引入基于VirusTotal时间戳的时间验证方法,评估推断出的树形结构是否反映真实的演化顺序。MalTree实现了87%的时间一致性,表明推断的演化关系与真实世界的出现时间线高度吻合。分析显示,某些家族的变异速度是其他家族的10倍以上,这表明检测策略应根据家族特异的演化节奏进行调整。包括Mirai僵尸网络在内的案例研究证实,从我们的系统发育树推断出的关系与已知威胁情报一致。该框架为将恶意软件分析从逐样本分类转向谱系感知的演化建模奠定了基础。