With the launch of ChatGPT, large language models (LLMs) have attracted global attention. In the realm of article writing, LLMs have witnessed extensive utilization, giving rise to concerns related to intellectual property protection, personal privacy, and academic integrity. In response, AI-text detection has emerged to distinguish between human and machine-generated content. However, recent research indicates that these detection systems often lack robustness and struggle to effectively differentiate perturbed texts. Currently, there is a lack of systematic evaluations regarding detection performance in real-world applications, and a comprehensive examination of perturbation techniques and detector robustness is also absent. To bridge this gap, our work simulates real-world scenarios in both informal and professional writing, exploring the out-of-the-box performance of current detectors. Additionally, we have constructed 12 black-box text perturbation methods to assess the robustness of current detection models across various perturbation granularities. Furthermore, through adversarial learning experiments, we investigate the impact of perturbation data augmentation on the robustness of AI-text detectors. We have released our code and data at https://github.com/zhouying20/ai-text-detector-evaluation.
翻译:随着ChatGPT的发布,大语言模型(LLMs)已引起全球关注。在文章写作领域,LLMs得到了广泛应用,同时也引发了关于知识产权保护、个人隐私和学术诚信的担忧。为此,AI文本检测技术应运而生,旨在区分人类与机器生成的内容。然而,近期研究表明,这些检测系统往往缺乏鲁棒性,难以有效识别经过扰动的文本。目前,对于实际应用场景中的检测性能尚缺乏系统性评估,同时也缺少对扰动技术与检测器鲁棒性的全面考察。为填补这一空白,本研究模拟了非正式与专业写作中的真实场景,探究了当前检测器在未经专门训练下的性能表现。此外,我们构建了12种黑盒文本扰动方法,以评估当前检测模型在不同扰动粒度下的鲁棒性。进一步地,通过对抗性学习实验,我们研究了扰动数据增强对AI文本检测器鲁棒性的影响。相关代码与数据已发布于https://github.com/zhouying20/ai-text-detector-evaluation。