The risk of language models unintentionally reproducing copyrighted material from their training data has led to the development of various protective measures. In this paper, we propose model fusion as an effective solution to safeguard against copyright infringement. In particular, we introduce Copyright-Protecting Fusion (CP-Fuse), an algorithm that adaptively combines language models to minimize the reproduction of protected materials. CP-Fuse is inspired by the recently proposed Near-Access Free (NAF) framework and additionally incorporates a desirable balancing property that we demonstrate prevents the reproduction of memorized training data. Our results show that CP-Fuse significantly reduces the memorization of copyrighted content while maintaining high-quality text and code generation. Furthermore, we demonstrate how CP-Fuse can be integrated with other techniques for enhanced protection.
翻译:语言模型无意中从其训练数据中复现受版权保护材料的风险,催生了各种保护措施的发展。本文提出模型融合作为一种防止版权侵权的有效解决方案。具体而言,我们引入了版权保护融合算法,该算法自适应地组合多个语言模型,以最小化受保护材料的复现。CP-Fuse的灵感来源于近期提出的近访问自由框架,并额外融入了一种理想的平衡特性;我们证明该特性能够防止已记忆训练数据的复现。我们的结果表明,CP-Fuse在保持高质量文本和代码生成能力的同时,显著降低了对受版权保护内容的记忆。此外,我们还展示了CP-Fuse如何与其他技术结合以实现增强保护。