The generations of large language models are commonly controlled through prompting techniques, where a user's query to the model is prefixed with a prompt that aims to guide the model's behaviour on the query. The prompts used by companies to guide their models are often treated as secrets, to be hidden from the user making the query. They have even been treated as commodities to be bought and sold. However, there has been anecdotal evidence showing that the prompts can be extracted by a user even when they are kept secret. In this paper, we present a framework for systematically measuring the success of prompt extraction attacks. In experiments with multiple sources of prompts and multiple underlying language models, we find that simple text-based attacks can in fact reveal prompts with high probability.
翻译:大型语言模型的生成通常通过提示技术来控制,其中用户的查询被附加上旨在引导模型在查询上行为的前缀提示。公司用于指导其模型的提示常被视为秘密,需对进行查询的用户隐藏,甚至被当作商品进行买卖。然而,有轶事证据表明,即使提示被保密,用户也能将其提取出来。在本文中,我们提出了一个框架,用于系统地衡量提示提取攻击的成功率。在涉及多个提示来源和多个底层语言模型的实验中,我们发现简单的基于文本的攻击实际上能以高概率揭示提示。