In this work, we study literature in Explainable AI and Safe AI to understand poisoning of neural models of code. In order to do so, we first establish a novel taxonomy for Trojan AI for code, and present a new aspect-based classification of triggers in neural models of code. Next, we highlight recent works that help us deepen our conception of how these models understand software code. Then we pick some of the recent, state-of-art poisoning strategies that can be used to manipulate such models. The insights we draw can potentially help to foster future research in the area of Trojan AI for code.
翻译:本文通过研究可解释人工智能与安全人工智能领域的文献,旨在理解神经代码模型的投毒机制。为此,我们首先建立了面向代码的神经木马新型分类体系,并提出了基于方面分类的代码神经模型触发器分类方法。随后,我们重点评述了近期有助于深化对模型理解软件代码机理的研究成果。在此基础上,我们选取了当前若干前沿的投毒策略,这些策略可用于操纵此类模型。本文所获得的见解有望推动神经代码木马领域的未来研究。