Large Language Models (LLMs) have achieved human-level fluency in text generation, making it difficult to distinguish between human-written and LLM-generated texts. This poses a growing risk of misuse of LLMs and demands the development of detectors to identify LLM-generated texts. However, existing detectors lack robustness against attacks: they degrade detection accuracy by simply paraphrasing LLM-generated texts. Furthermore, a malicious user might attempt to deliberately evade the detectors based on detection results, but this has not been assumed in previous studies. In this paper, we propose OUTFOX, a framework that improves the robustness of LLM-generated-text detectors by allowing both the detector and the attacker to consider each other's output. In this framework, the attacker uses the detector's prediction labels as examples for in-context learning and adversarially generates essays that are harder to detect, while the detector uses the adversarially generated essays as examples for in-context learning to learn to detect essays from a strong attacker. Experiments in the domain of student essays show that the proposed detector improves the detection performance on the attacker-generated texts by up to +41.3 points in F1-score. Furthermore, the proposed detector shows a state-of-the-art detection performance: up to 96.9 points in F1-score, beating existing detectors on non-attacked texts. Finally, the proposed attacker drastically degrades the performance of detectors by up to -57.0 points F1-score, massively outperforming the baseline paraphrasing method for evading detection.
翻译:摘要:大型语言模型(LLMs)在文本生成方面已达到人类水平的流畅度,使得区分人类撰写文本与LLM生成文本变得困难。这带来了LLM被滥用的日益增长的风险,从而亟需开发检测器识别LLM生成的文本。然而,现有检测器缺乏对抗攻击的鲁棒性:仅通过改写LLM生成文本就能降低其检测准确性。此外,恶意用户可能试图根据检测结果故意规避检测器,但这一点在先前研究中未被考虑。本文提出OUTFOX框架,通过允许检测器和攻击者相互考虑对方输出来提升LLM生成文本检测器的鲁棒性。在该框架中,攻击者利用检测器的预测标签作为上下文学习的示例,对抗性地生成更难被检测的论文;而检测器则使用这些对抗生成示例进行上下文学习,从而学会检测来自强力攻击者的论文。学生论文领域的实验表明,所提检测器在攻击者生成文本上的F1分数最高提升+41.3个点。此外,该检测器展现了最先进的检测性能:在未受攻击文本上的F1分数高达96.9个点,超越现有检测器。最后,所提攻击者使检测器性能大幅下降,F1分数最高降低-57.0个点,在规避检测方面显著优于作为基线的改写方法。