Diffusion models have achieved remarkable success in generating high-quality images thanks to their novel training procedures applied to unprecedented amounts of data. However, training a diffusion model from scratch is computationally expensive. This highlights the need to investigate the possibility of training these models iteratively, reusing computation while the data distribution changes. In this study, we take the first step in this direction and evaluate the continual learning (CL) properties of diffusion models. We begin by benchmarking the most common CL methods applied to Denoising Diffusion Probabilistic Models (DDPMs), where we note the strong performance of the experience replay with the reduced rehearsal coefficient. Furthermore, we provide insights into the dynamics of forgetting, which exhibit diverse behavior across diffusion timesteps. We also uncover certain pitfalls of using the bits-per-dimension metric for evaluating CL.
翻译:扩散模型凭借其应用于前所未有数据量的新型训练流程,在生成高质量图像方面取得了显著成功。然而,从头训练扩散模型的计算成本高昂。这凸显出有必要研究在数据分布变化时,通过迭代训练这些模型并重用计算资源的可能性。在本研究中,我们朝这一方向迈出了第一步,评估了扩散模型的持续学习特性。我们首先对应用于去噪扩散概率模型的最常见持续学习方法进行基准测试,发现采用缩减回放系数的经验回放方法表现强劲。此外,我们揭示了遗忘动态的见解,这些动态在扩散时间步长上表现出多样化的行为。我们还发现了使用每维比特数这一指标评估持续学习时的某些陷阱。