Large Language Models (LLMs) are known to memorize significant portions of their training data. Parts of this memorized content have been shown to be extractable by simply querying the model, which poses a privacy risk. We present a novel approach which uses prompt-tuning to control the extraction rates of memorized content in LLMs. We present two prompt training strategies to increase and decrease extraction rates, which correspond to an attack and a defense, respectively. We demonstrate the effectiveness of our techniques by using models from the GPT-Neo family on a public benchmark. For the 1.3B parameter GPT-Neo model, our attack yields a 9.3 percentage point increase in extraction rate compared to our baseline. Our defense can be tuned to achieve different privacy-utility trade-offs by a user-specified hyperparameter. We achieve an extraction rate reduction of up to 97.7% relative to our baseline, with a perplexity increase of 16.9%.
翻译:大型语言模型(LLMs)已被证实会记忆其训练数据的大部分内容。研究表明,仅通过查询模型即可提取部分记忆化内容,这构成了隐私风险。我们提出了一种基于提示调优(prompt-tuning)的新方法,用于控制LLMs中记忆化内容的提取率。我们设计了两种提示训练策略,分别用于提高和降低提取率,对应攻击与防御场景。通过在GPT-Neo系列模型上使用公开基准测试,我们验证了该技术的有效性。对于1.3B参数的GPT-Neo模型,我们的攻击方法相比基线将提取率提升了9.3个百分点。防御方法可通过用户指定的超参数调整隐私-效用的平衡,最终实现提取率相对基线降低97.7%,同时困惑度仅增加16.9%。