Vector quantized diffusion (VQ-Diffusion) is a powerful generative model for text-to-image synthesis, but sometimes can still generate low-quality samples or weakly correlated images with text input. We find these issues are mainly due to the flawed sampling strategy. In this paper, we propose two important techniques to further improve the sample quality of VQ-Diffusion. 1) We explore classifier-free guidance sampling for discrete denoising diffusion model and propose a more general and effective implementation of classifier-free guidance. 2) We present a high-quality inference strategy to alleviate the joint distribution issue in VQ-Diffusion. Finally, we conduct experiments on various datasets to validate their effectiveness and show that the improved VQ-Diffusion suppresses the vanilla version by large margins. We achieve an 8.44 FID score on MSCOCO, surpassing VQ-Diffusion by 5.42 FID score. When trained on ImageNet, we dramatically improve the FID score from 11.89 to 4.83, demonstrating the superiority of our proposed techniques.
翻译:向量量化扩散(VQ-Diffusion)是文本到图像合成领域强大的生成模型,但有时仍会产生低质量样本或与文本输入弱相关的图像。我们发现这些问题主要源于存在缺陷的采样策略。本文提出两项关键技术以进一步提升VQ-Diffusion的样本质量:1)我们探索了离散去噪扩散模型的无分类器引导采样,并提出更通用且高效的无分类器引导实现方法;2)提出高质量推理策略以缓解VQ-Diffusion中的联合分布问题。最后,我们在多个数据集上开展实验验证其有效性,结果表明改进版VQ-Diffusion在各项指标上显著超越原始版本。改进模型在MSCOCO数据集上取得8.44的FID分数,相较VQ-Diffusion提升5.42;在ImageNet上训练时,FID分数从11.89大幅降至4.83,充分证明了所提技术的优越性。