Testing of Detection Tools for AI-Generated Text

Recent advances in generative pre-trained transformer large language models have emphasised the potential risks of unfair use of artificial intelligence (AI) generated content in an academic environment and intensified efforts in searching for solutions to detect such content. The paper examines the general functionality of detection tools for artificial intelligence generated text and evaluates them based on accuracy and error type analysis. Specifically, the study seeks to answer research questions about whether existing detection tools can reliably differentiate between human-written text and ChatGPT-generated text, and whether machine translation and content obfuscation techniques affect the detection of AI-generated text. The research covers 12 publicly available tools and two commercial systems (Turnitin and PlagiarismCheck) that are widely used in the academic setting. The researchers conclude that the available detection tools are neither accurate nor reliable and have a main bias towards classifying the output as human-written rather than detecting AI-generated text. Furthermore, content obfuscation techniques significantly worsen the performance of tools. The study makes several significant contributions. First, it summarises up-to-date similar scientific and non-scientific efforts in the field. Second, it presents the result of one of the most comprehensive tests conducted so far, based on a rigorous research methodology, an original document set, and a broad coverage of tools. Third, it discusses the implications and drawbacks of using detection tools for AI-generated text in academic settings.

翻译：近期，生成式预训练Transformer大语言模型的进展凸显了在学术环境中不公平使用人工智能生成内容的潜在风险，并促使学界加大力度探索检测此类内容的方法。本文旨在分析人工智能生成文本检测工具的基本功能，并基于准确度与错误类型分析对其进行评估。具体而言，本研究尝试回答以下研究问题：现有检测工具能否可靠区分人类撰写的文本与ChatGPT生成的文本？机器翻译及内容混淆技术是否会影响AI生成文本的检测效果？研究涵盖了12款公开可用的检测工具及两种在学术领域广泛应用的商业系统（Turnitin和PlagiarismCheck）。研究人员得出结论：现有检测工具既不够精确也不可靠，且存在将输出结果归类为人类撰写而非AI生成文本的主要偏差。此外，内容混淆技术显著降低了工具的性能。本研究具有多项重要贡献：第一，系统总结了该领域最新的科学与非科学相关研究进展；第二，基于严谨的研究方法、原创文档集及广泛的工具覆盖范围，呈现了迄今为止最全面的检测测试结果之一；第三，探讨了在学术环境中使用AI生成文本检测工具的意义与局限性。

相关内容

TOOLS

关注 1

这个新版本的工具会议系列恢复了从1989年到2012年的50个会议的传统。工具最初是“面向对象语言和系统的技术”，后来发展到包括软件技术的所有创新方面。今天许多最重要的软件概念都是在这里首次引入的。2019年TOOLS 50+1在俄罗斯喀山附近举行，以同样的创新精神、对所有与软件相关的事物的热情、科学稳健性和行业适用性的结合以及欢迎该领域所有趋势和社区的开放态度，延续了该系列。官网链接：http://tools2019.innopolis.ru/

【CVPR 2022】一个完全无监督的框架，从噪声和部分测量中学习图像，Robust Equivariant Imaging: a fully unsupervised framework for learning to image

专知会员服务

25+阅读 · 2022年3月3日

O’Reilly报告：知识图谱崛起——面向现代数据集成和数据结构体系，“The Rise of the Knowledge Graph——Toward Modern Data Integration and the Data Fabric Architecture”

专知会员服务

49+阅读 · 2022年2月18日

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

35+阅读 · 2019年10月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日