Deep generative models have made tremendous progress in modeling complex data, often exhibiting generation quality that surpasses a typical human's ability to discern the authenticity of samples. Undeniably, a key driver of this success is enabled by the massive amounts of web-scale data consumed by these models. Due to these models' striking performance and ease of availability, the web will inevitably be increasingly populated with synthetic content. Such a fact directly implies that future iterations of generative models will be trained on both clean and artificially generated data from past models. In this paper, we develop a framework to rigorously study the impact of training generative models on mixed datasets -- from classical training on real data to self-consuming generative models trained on purely synthetic data. We first prove the stability of iterative training under the condition that the initial generative models approximate the data distribution well enough and the proportion of clean training data (w.r.t. synthetic data) is large enough. We empirically validate our theory on both synthetic and natural images by iteratively training normalizing flows and state-of-the-art diffusion models on CIFAR10 and FFHQ.
翻译:深度生成模型在建模复杂数据方面取得了巨大进展,常常展现出超越普通人辨别样本真实性能力的生成质量。不可否认,这一成功的关键驱动力源于这些模型所消耗的海量网络规模数据。由于这些模型卓越的性能和易获取性,网络将不可避免地日益充斥着合成内容。这一事实直接意味着未来的生成模型迭代版本将同时使用真实数据与先前模型生成的人工数据进行训练。本文建立了一个框架,严格研究在混合数据集上训练生成模型的影响——从经典的真实数据训练到完全基于合成数据的自消耗生成模型训练。我们首先证明:在初始生成模型足够好地逼近数据分布,且干净训练数据占比(相对于合成数据)足够大的条件下,迭代训练具有稳定性。通过在CIFAR10和FFHQ数据集上对归一化流和当前最优扩散模型进行迭代训练,我们在合成图像和自然图像上实证验证了我们的理论。