The rapid progress of Large Language Models (LLMs) has made them capable of performing astonishingly well on various tasks including document completion and question answering. The unregulated use of these models, however, can potentially lead to malicious consequences such as plagiarism, generating fake news, spamming, etc. Therefore, reliable detection of AI-generated text can be critical to ensure the responsible use of LLMs. Recent works attempt to tackle this problem either using certain model signatures present in the generated text outputs or by applying watermarking techniques that imprint specific patterns onto them. In this paper, both empirically and theoretically, we show that these detectors are not reliable in practical scenarios. Empirically, we show that paraphrasing attacks, where a light paraphraser is applied on top of the generative text model, can break a whole range of detectors, including the ones using the watermarking schemes as well as neural network-based detectors and zero-shot classifiers. We then provide a theoretical impossibility result indicating that for a sufficiently good language model, even the best-possible detector can only perform marginally better than a random classifier. Finally, we show that even LLMs protected by watermarking schemes can be vulnerable against spoofing attacks where adversarial humans can infer hidden watermarking signatures and add them to their generated text to be detected as text generated by the LLMs, potentially causing reputational damages to their developers. We believe these results can open an honest conversation in the community regarding the ethical and reliable use of AI-generated text.
翻译:大型语言模型(LLMs)的快速发展使其在文档补全、问答等多种任务中展现出惊人的性能。然而,这些模型的不受约束使用可能引发剽窃、生成虚假新闻、垃圾信息等恶意后果。因此,对AI生成文本的可靠检测对于确保LLMs的负责任使用至关重要。近期研究试图通过利用生成文本输出中的特定模型特征,或采用水印技术在其中嵌入特定模式来解决该问题。本文从实证与理论两方面表明,这些检测器在实际场景中并不可靠。实证上,我们证明通过轻量级释义攻击(即在生成文本模型基础上应用简易释义器)可突破包括采用水印方案、基于神经网络的检测器及零样本分类器在内的多种检测系统。随后我们给出理论上的不可能性结果:对于足够优质的语言模型,即便最优检测器也只能比随机分类器略优。最后我们揭示,即便采用水印保护的LLMs也易受欺骗攻击——对抗性人类可推断隐藏的水印特征并将其添加至自身生成文本中,使该文本被误判为LLMs生成内容,从而可能损害开发者的声誉。我们相信这些结果将推动学界就AI生成文本的伦理与可靠使用展开坦诚对话。