Though many deep learning-based models have made great progress in vulnerability detection, we have no good understanding of these models, which limits the further advancement of model capability, understanding of the mechanism of model detection, and efficiency and safety of practical application of models. In this paper, we extensively and comprehensively investigate two types of state-of-the-art learning-based approaches (sequence-based and graph-based) by conducting experiments on a recently built large-scale dataset. We investigate seven research questions from five dimensions, namely model capabilities, model interpretation, model stability, ease of use of model, and model economy. We experimentally demonstrate the priority of sequence-based models and the limited abilities of both LLM (ChatGPT) and graph-based models. We explore the types of vulnerability that learning-based models skilled in and reveal the instability of the models though the input is subtlely semantical-equivalently changed. We empirically explain what the models have learned. We summarize the pre-processing as well as requirements for easily using the models. Finally, we initially induce the vital information for economically and safely practical usage of these models.
翻译:尽管许多基于深度学习的模型在漏洞检测方面取得了巨大进展,但我们对这些模型缺乏深入理解,这限制了模型能力的进一步提升、模型检测机制的理解以及模型实际应用的效率与安全性。本文通过在近期构建的大规模数据集上进行实验,广泛而全面地研究了两种最先进的基于学习的方法(基于序列的方法和基于图的方法)。我们从五个维度探讨了七个研究问题,即模型能力、模型可解释性、模型稳定性、模型易用性以及模型经济性。实验结果表明,基于序列的模型具有优势,而大型语言模型(ChatGPT)和基于图的模型能力有限。我们探究了基于学习的模型擅长检测的漏洞类型,并揭示了即使输入发生细微的语义等价变化,模型也会表现出不稳定性。我们通过实证解释了模型学习到的内容。总结了模型的预处理步骤及易用性要求。最后,我们初步归纳了经济且安全地实际应用这些模型所需的关键信息。