There is a growing concern that learned conditional generative models may output samples that are substantially similar to some copyrighted data $C$ that was in their training set. We give a formal definition of $\textit{near access-freeness (NAF)}$ and prove bounds on the probability that a model satisfying this definition outputs a sample similar to $C$, even if $C$ is included in its training set. Roughly speaking, a generative model $p$ is $\textit{$k$-NAF}$ if for every potentially copyrighted data $C$, the output of $p$ diverges by at most $k$-bits from the output of a model $q$ that $\textit{did not access $C$ at all}$. We also give generative model learning algorithms, which efficiently modify the original generative model learning algorithm in a black box manner, that output generative models with strong bounds on the probability of sampling protected content. Furthermore, we provide promising experiments for both language (transformers) and image (diffusion) generative models, showing minimal degradation in output quality while ensuring strong protections against sampling protected content.
翻译:人们日益关注到,学习到的条件生成模型可能会输出与其训练集中某些受版权保护数据 $C$ 高度相似的样本。我们给出了$\textit{近似无访问性(NAF)}$的形式化定义,并证明了满足该定义的模型输出与 $C$ 相似样本的概率上界,即使 $C$ 包含在其训练集中。粗略而言,若对于任意可能受版权保护的数据 $C$,生成模型 $p$ 的输出与一个$\textit{从未访问过 $C$}$的模型 $q$ 的输出最多相差 $k$ 比特,则称 $p$ 为$\textit{$k$-NAF}$。我们还提出了生成模型学习算法,该算法以黑盒方式高效修改原始生成模型学习算法,输出具有强概率上界(保护受保护内容不被采样)的生成模型。此外,我们为语言(Transformer)和图像(扩散)生成模型提供了有前景的实验结果,表明在确保对受保护内容采样的强保护的同时,输出质量退化极小。