There is a lack of research into capabilities of recent LLMs to generate convincing text in languages other than English and into performance of detectors of machine-generated text in multilingual settings. This is also reflected in the available benchmarks which lack authentic texts in languages other than English and predominantly cover older generators. To fill this gap, we introduce MULTITuDE, a novel benchmarking dataset for multilingual machine-generated text detection comprising of 74,081 authentic and machine-generated texts in 11 languages (ar, ca, cs, de, en, es, nl, pt, ru, uk, and zh) generated by 8 multilingual LLMs. Using this benchmark, we compare the performance of zero-shot (statistical and black-box) and fine-tuned detectors. Considering the multilinguality, we evaluate 1) how these detectors generalize to unseen languages (linguistically similar as well as dissimilar) and unseen LLMs and 2) whether the detectors improve their performance when trained on multiple languages.
翻译:当前研究在以下方面存在不足:评估最新大语言模型(LLM)生成非英语语言可信文本的能力,以及多语言环境下机器生成文本检测器的性能。这一问题同样反映在现有基准数据集中——它们缺乏非英语语言的真实文本,且主要覆盖早期生成器。为填补这一空白,我们提出MULTITuDE——一个面向多语言机器生成文本检测的新型基准数据集,包含由8个多语言LLM生成的74,081篇真实与机器生成文本,覆盖11种语言(ar、ca、cs、de、en、es、nl、pt、ru、uk、zh)。基于该基准,我们比较了零样本(统计与黑箱)检测器与微调检测器的性能。针对多语言特性,我们评估了:1)检测器如何泛化至未见语言(包括语言相似与不相似的情况)及未见LLM;2)在多种语言上训练是否能够提升检测器性能。