Machine generated text is increasingly difficult to distinguish from human authored text. Powerful open-source models are freely available, and user-friendly tools that democratize access to generative models are proliferating. ChatGPT, which was released shortly after the first preprint of this survey, epitomizes these trends. The great potential of state-of-the-art natural language generation (NLG) systems is tempered by the multitude of avenues for abuse. Detection of machine generated text is a key countermeasure for reducing abuse of NLG models, with significant technical challenges and numerous open problems. We provide a survey that includes both 1) an extensive analysis of threat models posed by contemporary NLG systems, and 2) the most complete review of machine generated text detection methods to date. This survey places machine generated text within its cybersecurity and social context, and provides strong guidance for future work addressing the most critical threat models, and ensuring detection systems themselves demonstrate trustworthiness through fairness, robustness, and accountability.
翻译:机器生成文本日益难以与人类撰写的文本区分。强大的开源模型可免费获取,而普及生成式模型访问权限的用户友好型工具正在迅速涌现。本综述初版预印本发布后不久面世的ChatGPT,正是这些趋势的典型代表。最先进的自然语言生成系统虽潜力巨大,但其滥用途径众多,限制了这些潜能的发挥。机器生成文本检测是减少NLG模型滥用的关键对策,但面临重大技术挑战和大量未解决问题。本综述包含两方面内容:1)对当代NLG系统构成的威胁模型进行全面分析;2)提供迄今为止最完整的机器生成文本检测方法综述。本综述将机器生成文本置于网络安全与社会背景中,为未来聚焦最关键威胁模型、通过公平性、鲁棒性和可问责性确保检测系统本身可信的研究提供了有力指导。