How do diffusion generative models convert pure noise into meaningful images? We argue that generation involves first committing to an outline, and then to finer and finer details. The corresponding reverse diffusion process can be modeled by dynamics on a (time-dependent) high-dimensional landscape full of Gaussian-like modes, which makes the following predictions: (i) individual trajectories tend to be very low-dimensional; (ii) scene elements that vary more within training data tend to emerge earlier; and (iii) early perturbations substantially change image content more often than late perturbations. We show that the behavior of a variety of trained unconditional and conditional diffusion models like Stable Diffusion is consistent with these predictions. Finally, we use our theory to search for the latent image manifold of diffusion models, and propose a new way to generate interpretable image variations. Our viewpoint suggests generation by GANs and diffusion models have unexpected similarities.
翻译:扩散生成模型如何将纯噪声转化为有意义的图像?我们认为生成过程首先确定轮廓,然后逐步细化细节。对应的逆向扩散过程可以用充满类高斯模态的(时间相关)高维景观上的动力学建模,这做出以下预测:(1)单个轨迹往往具有极低维度;(2)训练数据中变化较大的场景元素倾向于更早出现;(3)早期扰动比后期扰动更频繁地显著改变图像内容。我们证明多种无条件和条件扩散模型(如Stable Diffusion)的行为与这些预测一致。最后,利用该理论探索扩散模型的潜在图像流形,并提出生成可解释图像变体的新方法。我们的观点表明,GAN与扩散模型的生成过程存在意想不到的相似性。