Authorship attribution has become increasingly accurate, posing a serious privacy risk for programmers who wish to remain anonymous. In this paper, we introduce SHIELD to examine the robustness of different code authorship attribution approaches against adversarial code examples. We define four attacks on attribution techniques, which include targeted and non-targeted attacks, and realize them using adversarial code perturbation. We experiment with a dataset of 200 programmers from the Google Code Jam competition to validate our methods targeting six state-of-the-art authorship attribution methods that adopt a variety of techniques for extracting authorship traits from source-code, including RNN, CNN, and code stylometry. Our experiments demonstrate the vulnerability of current authorship attribution methods against adversarial attacks. For the non-targeted attack, our experiments demonstrate the vulnerability of current authorship attribution methods against the attack with an attack success rate exceeds 98.5\% accompanied by a degradation of the identification confidence that exceeds 13\%. For the targeted attacks, we show the possibility of impersonating a programmer using targeted-adversarial perturbations with a success rate ranging from 66\% to 88\% for different authorship attribution techniques under several adversarial scenarios.
翻译:作者归属识别技术日益精准,对希望保持匿名的程序员构成了严重的隐私风险。本文提出SHIELD方法,旨在检验不同代码作者归属方法面对对抗性代码样本时的鲁棒性。我们定义了针对归属技术的四种攻击方式,包括定向与非定向攻击,并通过对抗性代码扰动实现这些攻击。基于Google Code Jam竞赛中200名程序员的实验数据集,我们验证了针对六种采用不同源码作者特征提取技术(包括RNN、CNN及代码风格计量学)的最新作者归属方法的效果。实验表明,当前作者归属方法在面对对抗性攻击时存在脆弱性。在非定向攻击中,我们的实验显示攻击成功率超过98.5%,同时识别置信度下降超过13%。在定向攻击中,我们证明通过定向对抗扰动可模拟特定程序员身份,在不同对抗场景下对不同作者归属技术的模仿成功率介于66%至88%之间。