Large language models (LLMs) have witnessed a meteoric rise in popularity among the general public users over the past few months, facilitating diverse downstream tasks with human-level accuracy and proficiency. Prompts play an essential role in this success, which efficiently adapt pre-trained LLMs to task-specific applications by simply prepending a sequence of tokens to the query texts. However, designing and selecting an optimal prompt can be both expensive and demanding, leading to the emergence of Prompt-as-a-Service providers who profit by providing well-designed prompts for authorized use. With the growing popularity of prompts and their indispensable role in LLM-based services, there is an urgent need to protect the copyright of prompts against unauthorized use. In this paper, we propose PromptCARE, the first framework for prompt copyright protection through watermark injection and verification. Prompt watermarking presents unique challenges that render existing watermarking techniques developed for model and dataset copyright verification ineffective. PromptCARE overcomes these hurdles by proposing watermark injection and verification schemes tailor-made for prompts and NLP characteristics. Extensive experiments on six well-known benchmark datasets, using three prevalent pre-trained LLMs (BERT, RoBERTa, and Facebook OPT-1.3b), demonstrate the effectiveness, harmlessness, robustness, and stealthiness of PromptCARE.
翻译:大型语言模型在近几个月间于普通用户群体中迅速普及,能以人类级别的准确性和熟练度支持多种下游任务。提示在此成功中扮演关键角色——通过在查询文本前预置标记序列,高效地将预训练大语言模型适配至特定任务应用。然而,设计并选择最优提示既昂贵又要求严苛,由此催生了提示即服务提供商,通过授权使用精心设计的提示获利。随着提示的广泛应用及其在大语言模型服务中的核心地位,亟需保护提示版权免遭未授权使用。本文提出PromptCARE,首个通过水印注入与验证实现提示版权保护的框架。提示水印面临的独特挑战,使现有面向模型与数据集版权验证的水印技术难以奏效。PromptCARE通过提出专为提示及自然语言处理特性设计的水印注入与验证方案克服这些障碍。基于三个主流预训练大语言模型在六个知名基准数据集上的广泛实验表明,PromptCARE具有有效性、无害性、鲁棒性与隐蔽性。