Nowadays, powerful large language models (LLMs) such as ChatGPT have demonstrated revolutionary power in a variety of tasks. Consequently, the detection of machine-generated texts (MGTs) is becoming increasingly crucial as LLMs become more advanced and prevalent. These models have the ability to generate human-like language, making it challenging to discern whether a text is authored by a human or a machine. This raises concerns regarding authenticity, accountability, and potential bias. However, existing methods for detecting MGTs are evaluated using different model architectures, datasets, and experimental settings, resulting in a lack of a comprehensive evaluation framework that encompasses various methodologies. Furthermore, it remains unclear how existing detection methods would perform against powerful LLMs. In this paper, we fill this gap by proposing the first benchmark framework for MGT detection against powerful LLMs, named MGTBench. Extensive evaluations on public datasets with curated texts generated by various powerful LLMs such as ChatGPT-turbo and Claude demonstrate the effectiveness of different detection methods. Our ablation study shows that a larger number of words in general leads to better performance and most detection methods can achieve similar performance with much fewer training samples. Moreover, we delve into a more challenging task: text attribution. Our findings indicate that the model-based detection methods still perform well in the text attribution task. To investigate the robustness of different detection methods, we consider three adversarial attacks, namely paraphrasing, random spacing, and adversarial perturbations. We discover that these attacks can significantly diminish detection effectiveness, underscoring the critical need for the development of more robust detection methods.
翻译:如今,诸如ChatGPT等强大的大型语言模型已在多种任务中展现出革命性能力。随着这些模型日益先进和普及,机器生成文本的检测变得愈加关键。这些模型能够生成类人语言,使得辨别文本由人类还是机器撰写变得极具挑战性,由此引发了对真实性、问责性和潜在偏差的担忧。然而,现有检测机器生成文本的方法采用不同的模型架构、数据集和实验设置进行评估,导致缺乏一个涵盖多种方法的综合评估框架。此外,现有检测方法在面对强大大型语言模型时的表现尚不明确。本文首次提出面向强大大型语言模型的机器生成文本检测基准框架MGTBench,填补了这一空白。通过在包含ChatGPT-turbo和Claude等多种强大大型语言模型生成文本的公开数据集上进行广泛评估,验证了不同检测方法的有效性。消融研究表明,更多单词数量通常能带来更优性能,且多数检测方法使用更少训练样本即可达到相近效果。此外,我们深入研究了更具挑战性的文本归因任务,结果表明基于模型的检测方法在该任务中仍表现良好。为探究不同检测方法的鲁棒性,我们考虑了三种对抗攻击——释义改写、随机间距和对抗扰动。实验发现这些攻击会显著降低检测效能,凸显了开发更鲁棒检测方法的迫切需求。