The rapid adoption of generative language models has brought about substantial advancements in digital communication, while simultaneously raising concerns regarding the potential misuse of AI-generated content. Although numerous detection methods have been proposed to differentiate between AI and human-generated content, the fairness and robustness of these detectors remain underexplored. In this study, we evaluate the performance of several widely-used GPT detectors using writing samples from native and non-native English writers. Our findings reveal that these detectors consistently misclassify non-native English writing samples as AI-generated, whereas native writing samples are accurately identified. Furthermore, we demonstrate that simple prompting strategies can not only mitigate this bias but also effectively bypass GPT detectors, suggesting that GPT detectors may unintentionally penalize writers with constrained linguistic expressions. Our results call for a broader conversation about the ethical implications of deploying ChatGPT content detectors and caution against their use in evaluative or educational settings, particularly when they may inadvertently penalize or exclude non-native English speakers from the global discourse. The published version of this study can be accessed at: www.cell.com/patterns/fulltext/S2666-3899(23)00130-7
翻译:生成式语言模型的快速普及在推动数字通信领域重大进展的同时,也引发了对AI生成内容潜在滥用现象的担忧。尽管已有大量检测方法被提出用于区分AI与人类生成内容,但这些检测器的公平性和鲁棒性仍待深入探究。本研究利用母语和非母语英语写作者的写作样本,评估了多种广泛应用的GPT检测器的性能。研究结果显示,这些检测器持续将非母语英语写作者的样本误判为AI生成,而母语写作者的样本则能被准确识别。此外,我们发现简单的提示策略不仅能够缓解这种偏见,还能有效绕过GPT检测器,表明GPT检测器可能无意中惩罚了语言表达能力受限的写作者。本研究结果呼吁对部署ChatGPT内容检测器的伦理影响展开更广泛的讨论,并警示在评估或教育场景中谨慎使用此类工具——特别是在可能无意中惩罚或排挤非母语英语使用者、使其难以参与全球对话的情境下。本研究已发表版本可访问:www.cell.com/patterns/fulltext/S2666-3899(23)00130-7