Classifier-free guided diffusion models have recently been shown to be highly effective at high-resolution image generation, and they have been widely used in large-scale diffusion frameworks including DALLE-2, Stable Diffusion and Imagen. However, a downside of classifier-free guided diffusion models is that they are computationally expensive at inference time since they require evaluating two diffusion models, a class-conditional model and an unconditional model, tens to hundreds of times. To deal with this limitation, we propose an approach to distilling classifier-free guided diffusion models into models that are fast to sample from: Given a pre-trained classifier-free guided model, we first learn a single model to match the output of the combined conditional and unconditional models, and then we progressively distill that model to a diffusion model that requires much fewer sampling steps. For standard diffusion models trained on the pixel-space, our approach is able to generate images visually comparable to that of the original model using as few as 4 sampling steps on ImageNet 64x64 and CIFAR-10, achieving FID/IS scores comparable to that of the original model while being up to 256 times faster to sample from. For diffusion models trained on the latent-space (e.g., Stable Diffusion), our approach is able to generate high-fidelity images using as few as 1 to 4 denoising steps, accelerating inference by at least 10-fold compared to existing methods on ImageNet 256x256 and LAION datasets. We further demonstrate the effectiveness of our approach on text-guided image editing and inpainting, where our distilled model is able to generate high-quality results using as few as 2-4 denoising steps.


翻译:无分类器引导的扩散模型最近被证明在高分辨率图像生成方面非常有效,并已广泛应用于包括DALLE-2、Stable Diffusion和Imagen在内的大规模扩散框架中。然而,无分类器引导扩散模型的一个缺点是推理时计算成本高昂,因为它需要评估两个扩散模型(一个类条件模型和一个无条件模型),重复数十到数百次。为应对这一限制,我们提出了一种将无分类器引导扩散模型蒸馏为快速采样模型的方法:给定一个预训练的无分类器引导模型,我们首先学习单个模型以匹配组合的条件和无条件模型的输出,然后逐步将该模型蒸馏为需要更少采样步骤的扩散模型。对于在像素空间上训练的标准扩散模型,我们的方法能够在ImageNet 64x64和CIFAR-10上仅用4个采样步骤生成与原始模型视觉上相当的图像,FID/IS分数与原始模型相当,同时采样速度提升高达256倍。对于在潜在空间上训练的扩散模型(例如Stable Diffusion),我们的方法能够在ImageNet 256x256和LAION数据集上仅用1到4个去噪步骤生成高保真图像,与现有方法相比,推理加速至少10倍。我们进一步在文本引导的图像编辑和修复任务中展示了该方法的有效性,其中蒸馏模型仅需2-4个去噪步骤即可生成高质量结果。

0
下载
关闭预览

相关内容

扩散模型是近年来快速发展并得到广泛关注的生成模型。它通过一系列的加噪和去噪过程,在复杂的图像分布和高斯分布之间建立联系,使得模型最终能将随机采样的高斯噪声逐步去噪得到一张图像。
【AAAI2023】用于复杂场景图像合成的特征金字塔扩散模型
视觉的有效扩散模型综述
专知会员服务
97+阅读 · 2022年10月20日
AAAI2022-无需蒸馏信号的对比学习小模型训练效能研究
专知会员服务
18+阅读 · 2021年12月23日
图卷积神经网络蒸馏知识,Distillating Knowledge from GCN
专知会员服务
96+阅读 · 2020年3月25日
浅谈扩散模型的有分类器引导和无分类器引导
PaperWeekly
4+阅读 · 2022年12月1日
扩散模型初探:原理及应用
PaperWeekly
3+阅读 · 2022年11月1日
国家自然科学基金
5+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
Arxiv
0+阅读 · 2023年5月31日
Arxiv
0+阅读 · 2023年5月31日
Arxiv
0+阅读 · 2023年5月30日
Arxiv
0+阅读 · 2023年5月29日
Arxiv
30+阅读 · 2022年9月10日
VIP会员
最新内容
乌克兰与中东为印太地区带来关于新太空战的启示
战争不仅需要机器人:人类仍不可或缺
专知会员服务
0+阅读 · 今天13:03
《描绘美国防部创新基础设施的未来蓝图》100页
专知会员服务
1+阅读 · 今天12:57
相关基金
国家自然科学基金
5+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员