There is a growing concern that learned conditional generative models may output samples that are substantially similar to some copyrighted data $C$ that was in their training set. We give a formal definition of $\textit{near access-freeness (NAF)}$ and prove bounds on the probability that a model satisfying this definition outputs a sample similar to $C$, even if $C$ is included in its training set. Roughly speaking, a generative model $p$ is $\textit{$k$-NAF}$ if for every potentially copyrighted data $C$, the output of $p$ diverges by at most $k$-bits from the output of a model $q$ that $\textit{did not access $C$ at all}$. We also give generative model learning algorithms, which efficiently modify the original generative model learning algorithm in a black box manner, that output generative models with strong bounds on the probability of sampling protected content. Furthermore, we provide promising experiments for both language (transformers) and image (diffusion) generative models, showing minimal degradation in output quality while ensuring strong protections against sampling protected content.
翻译:学界日益关注经学习的条件生成模型可能输出与训练集中受版权保护数据$C$高度相似的样本。本文给出$\textit{近无接触性(NAF)}$的形式化定义,并证明满足该定义的模型输出与$C$相似样本的概率上界——即使$C$包含在训练集中。直观而言,若生成模型$p$对任意潜在受版权保护数据$C$的输出,与$\textit{完全未接触}$数据$C$的模型$q$的输出差异不超过$k$比特,则称$p$为$\textit{$k$-近无接触}$。我们同时提出生成模型学习算法,通过黑盒方式高效修改原始生成模型学习过程,使输出生成模型具有对受保护内容采样概率的强约束。此外,针对语言(Transformer)和图像(扩散)生成模型开展的前瞻性实验表明,该方法在确保采样受保护内容的强防护能力的同时,模型输出质量仅有极小损失。