Since ChatGPT was introduced in November 2022, embedding (nearly) unnoticeable statistical signals into text generated by large language models (LLMs), also known as watermarking, has been used as a principled approach to provable detection of LLM-generated text from its human-written counterpart. In this paper, we introduce a general and flexible framework for reasoning about the statistical efficiency of watermarks and designing powerful detection rules. Inspired by the hypothesis testing formulation of watermark detection, our framework starts by selecting a pivotal statistic of the text and a secret key -- provided by the LLM to the verifier -- to enable controlling the false positive rate (the error of mistakenly detecting human-written text as LLM-generated). Next, this framework allows one to evaluate the power of watermark detection rules by obtaining a closed-form expression of the asymptotic false negative rate (the error of incorrectly classifying LLM-generated text as human-written). Our framework further reduces the problem of determining the optimal detection rule to solving a minimax optimization program. We apply this framework to two representative watermarks -- one of which has been internally implemented at OpenAI -- and obtain several findings that can be instrumental in guiding the practice of implementing watermarks. In particular, we derive optimal detection rules for these watermarks under our framework. These theoretically derived detection rules are demonstrated to be competitive and sometimes enjoy a higher power than existing detection approaches through numerical experiments.
翻译:自2022年11月ChatGPT问世以来,在大型语言模型(LLM)生成的文本中嵌入(近乎)不可察觉的统计信号(即水印技术)已被用作一种基于原理的方法,用于可证明地检测LLM生成的文本与人类撰写文本之间的区别。本文引入了一个通用且灵活的框架,用于推理水印的统计效率并设计强大的检测规则。受水印检测的假设检验公式启发,我们的框架首先选择文本的枢轴统计量和密钥——由LLM提供给验证者——以控制假阳性率(将人类撰写文本误检为LLM生成文本的错误)。接着,该框架通过获取渐近假阴性率(将LLM生成文本误分类为人类撰写文本的错误)的闭式表达式,允许评估水印检测规则的效力。我们的框架进一步将确定最优检测规则的问题简化为求解极小极大优化程序。我们将该框架应用于两个代表性的水印——其中一种已在OpenAI内部实现——并获得若干对指导水印实践具有重要意义的发现。特别地,我们在该框架下推导了这些水印的最优检测规则。数值实验表明,这些理论推导的检测规则具有竞争力,且在某些情况下比现有检测方法享有更高的效力。