Large Language models (LLMs) are trained on large amounts of data, which can include sensitive information that may compromise personal privacy. LLMs showed to memorize parts of the training data and emit those data verbatim when an adversary prompts appropriately. Previous research has primarily focused on data preprocessing and differential privacy techniques to address memorization or prevent verbatim memorization exclusively, which can give a false sense of privacy. However, these methods rely on explicit and implicit assumptions about the structure of the data to be protected, which often results in an incomplete solution to the problem. To address this, we propose a novel framework that utilizes a reinforcement learning approach (PPO) to fine-tune LLMs to mitigate approximate memorization. Our approach utilizes a negative similarity score, such as BERTScore or SacreBLEU, as a reward signal to learn a dissimilarity policy. Our results demonstrate that this framework effectively mitigates approximate memorization while maintaining high levels of coherence and fluency in the generated samples. Furthermore, our framework is robust in mitigating approximate memorization across various circumstances, including longer context, which is known to increase memorization in LLMs.
翻译:大语言模型(LLMs)在大量数据上进行训练,这些数据可能包含损害个人隐私的敏感信息。研究表明,LLMs能够记忆训练数据的部分内容,并在对手进行适当提示时逐字输出这些数据。以往的研究主要集中于数据预处理和差分隐私技术,以解决记忆问题或仅防止逐字记忆,这可能导致对隐私的虚假安全感。然而,这些方法依赖于对受保护数据结构的外显或内隐假设,往往导致问题的不完全解决。为此,我们提出一种新颖框架,利用强化学习方法(PPO)对LLMs进行微调,以缓解近似记忆。我们的方法采用负相似度得分(如BERTScore或SacreBLEU)作为奖励信号,学习相异策略。结果表明,该框架能有效缓解近似记忆,同时保持生成样本的高连贯性和流畅性。此外,该框架在多种情境下(包括已知会增加LLMs记忆的长上下文场景)均表现出对缓解近似记忆的稳健性。