We introduce Ghostbuster, a state-of-the-art system for detecting AI-generated text. Our method works by passing documents through a series of weaker language models, running a structured search over possible combinations of their features, and then training a classifier on the selected features to predict whether documents are AI-generated. Crucially, Ghostbuster does not require access to token probabilities from the target model, making it useful for detecting text generated by black-box models or unknown model versions. In conjunction with our model, we release three new datasets of human- and AI-generated text as detection benchmarks in the domains of student essays, creative writing, and news articles. We compare Ghostbuster to a variety of existing detectors, including DetectGPT and GPTZero, as well as a new RoBERTa baseline. Ghostbuster achieves 99.0 F1 when evaluated across domains, which is 5.9 F1 higher than the best preexisting model. It also outperforms all previous approaches in generalization across writing domains (+7.5 F1), prompting strategies (+2.1 F1), and language models (+4.4 F1). We also analyze the robustness of our system to a variety of perturbations and paraphrasing attacks and evaluate its performance on documents written by non-native English speakers.
翻译:我们提出了Ghostbuster——一种用于检测AI生成文本的前沿系统。该方法通过将文档输入一系列较弱的语言模型,对其特征的潜在组合进行结构化搜索,并在所选特征上训练分类器来预测文档是否为AI生成。关键之处在于,Ghostbuster无需访问目标模型的词元概率,因此适用于检测黑盒模型或未知模型版本生成的文本。结合该模型,我们发布了三个新的人机文本数据集,分别涵盖学生论文、创意写作和新闻文章领域,作为检测基准。我们将Ghostbuster与现有检测器(包括DetectGPT、GPTZero及新的RoBERTa基线模型)进行对比。跨领域评估显示,Ghostbuster的F1值达到99.0,比最优现有模型高出5.9 F1。在写作领域泛化(+7.5 F1)、提示策略泛化(+2.1 F1)和语言模型泛化(+4.4 F1)方面,该系统均优于以往所有方法。我们还分析了系统对各种干扰和改写攻击的鲁棒性,并评估了其在非英语母语者撰写文档上的表现。