Artificial Intelligence Generated Content (AIGC) has garnered considerable attention for its impressive performance, with ChatGPT emerging as a leading AIGC model that produces high-quality responses across various applications, including software development and maintenance. Despite its potential, the misuse of ChatGPT poses significant concerns, especially in education and safetycritical domains. Numerous AIGC detectors have been developed and evaluated on natural language data. However, their performance on code-related content generated by ChatGPT remains unexplored. To fill this gap, in this paper, we present the first empirical study on evaluating existing AIGC detectors in the software domain. We created a comprehensive dataset including 492.5K samples comprising code-related content produced by ChatGPT, encompassing popular software activities like Q&A (115K), code summarization (126K), and code generation (226.5K). We evaluated six AIGC detectors, including three commercial and three open-source solutions, assessing their performance on this dataset. Additionally, we conducted a human study to understand human detection capabilities and compare them with the existing AIGC detectors. Our results indicate that AIGC detectors demonstrate lower performance on code-related data compared to natural language data. Fine-tuning can enhance detector performance, especially for content within the same domain; but generalization remains a challenge. The human evaluation reveals that detection by humans is quite challenging.
翻译:人工智能生成内容(AIGC)因其卓越性能而受到广泛关注,其中ChatGPT作为领先的AIGC模型,在软件开发与维护等多种应用中生成高质量回复。尽管潜力巨大,ChatGPT的滥用引发了显著担忧,尤其在教育和安全关键领域。已有许多AIGC检测器在自然语言数据上得到开发与评估,然而它们对ChatGPT生成的代码相关内容的性能尚待探究。为填补这一空白,本文首次对软件领域现有AIGC检测器进行实证研究。我们构建了一个包含492.5K样本的综合数据集,涵盖ChatGPT生成的代码相关内容,涉及问答(115K)、代码摘要(126K)和代码生成(226.5K)等主流软件活动。我们评估了六种AIGC检测器(包括三种商业和三种开源解决方案),并测试它们在该数据集上的表现。此外,我们通过人工研究了解人类检测能力,并与现有AIGC检测器进行对比。结果表明,相较于自然语言数据,AIGC检测器在代码相关数据上的表现较低。微调可提升检测器性能,尤其对于同域内容,但泛化仍具挑战。人工评估揭示,人类检测相当困难。