Software vulnerability detection is critical for ensuring software security and reliability. Despite recent advances in deep learning, real-world vulnerability datasets suffer from two severe challenges: frequency imbalance and difficulty imbalance. We reinterpret these challenges from an embedding geometry perspective, observing that such imbalances induce geometric distortions in hyperspherical representation space. To address this issue, we propose MARGIN, a metric-based framework that learns discriminative vulnerability representations through adaptive margin metric learning and hyperspherical prototype modeling. MARGIN dynamically adjusts geometric regularization according to the distribution structure estimated by the von Mises-Fisher concentration, aligning the probability mass of embedding distributions with their corresponding Voronoi cells, thereby reducing geometric distortion and yielding more stable decision boundaries. Extensive experiments on public vulnerability datasets show that MARGIN consistently outperforms strong baselines, achieving notable improvements in classification and detection, especially on challenging, imbalanced datasets. Further analysis demonstrates that MARGIN produces more structured embedding geometries, improving robustness, interpretability, and generalization.
翻译:软件漏洞检测对保障软件安全性与可靠性至关重要。尽管深度学习近期取得显著进展,现实世界的漏洞数据集仍面临两大严峻挑战:频率不平衡与难度不平衡。我们从嵌入几何角度重新解读这些挑战,观察到上述不平衡会导致超球面表示空间产生几何畸变。为解决该问题,我们提出MARGIN——一种基于度量的框架,通过自适应边界度量学习与超球面原型建模来学习具有判别性的漏洞表示。MARGIN根据冯·米塞斯-费希尔浓度估计的分布结构动态调整几何正则化,将嵌入分布的概率质量对齐至对应的沃罗诺伊单元,从而减小几何畸变并生成更稳定的决策边界。在公开漏洞数据集上的大量实验表明,MARGIN持续优于强基线方法,在分类与检测任务中实现显著性能提升,尤其在具有挑战性的不平衡数据集上表现突出。进一步分析证实,MARGIN能生成更结构化的嵌入几何形态,从而增强鲁棒性、可解释性与泛化能力。