Several advances in deep learning have been successfully applied to the software development process. Of recent interest is the use of neural language models to build tools, such as Copilot, that assist in writing code. In this paper we perform a comparative empirical analysis of Copilot-generated code from a security perspective. The aim of this study is to determine if Copilot is as bad as human developers. We investigate whether Copilot is just as likely to introduce the same software vulnerabilities as human developers. Using a dataset of C/C++ vulnerabilities, we prompt Copilot to generate suggestions in scenarios that led to the introduction of vulnerabilities by human developers. The suggestions are inspected and categorized in a 2-stage process based on whether the original vulnerability or fix is reintroduced. We find that Copilot replicates the original vulnerable code about 33% of the time while replicating the fixed code at a 25% rate. However this behaviour is not consistent: Copilot is more likely to introduce some types of vulnerabilities than others and is also more likely to generate vulnerable code in response to prompts that correspond to older vulnerabilities. Overall, given that in a significant number of cases it did not replicate the vulnerabilities previously introduced by human developers, we conclude that Copilot, despite performing differently across various vulnerability types, is not as bad as human developers at introducing vulnerabilities in code.
翻译:深度学习领域的多项进展已成功应用于软件开发过程。近期备受关注的是利用神经语言模型构建诸如Copilot等辅助编写代码的工具。本文从安全角度对Copilot生成的代码进行了比较性实证分析,旨在探究Copilot是否与人类开发者同样糟糕。我们研究了Copilot是否与人类开发者一样容易引入相同的软件漏洞。利用C/C++漏洞数据集,我们促使Copilot在曾导致人类开发者引入漏洞的场景中生成建议。通过两阶段流程对这些建议进行审查和分类,判断原始漏洞或修复是否再次出现。结果发现,Copilot约有33%的概率复现原始漏洞代码,而以25%的概率复现修复代码。然而这种行为并不一致:Copilot更易引入某些类型的漏洞,且对对应较旧漏洞的提示更可能生成脆弱代码。总体而言,考虑到在大量案例中它并未复现人类开发者先前引入的漏洞,我们得出结论:尽管在不同漏洞类型上表现各异,Copilot在代码中引入漏洞的程度上并不比人类开发者更糟糕。