The recent wave of large-scale text-to-image diffusion models has dramatically increased our text-based image generation abilities. These models can generate realistic images for a staggering variety of prompts and exhibit impressive compositional generalization abilities. Almost all use cases thus far have solely focused on sampling; however, diffusion models can also provide conditional density estimates, which are useful for tasks beyond image generation. In this paper, we show that the density estimates from large-scale text-to-image diffusion models like Stable Diffusion can be leveraged to perform zero-shot classification without any additional training. Our generative approach to classification attains strong results on a variety of benchmarks and outperforms alternative methods of extracting knowledge from diffusion models. We also find that our diffusion-based approach has stronger multimodal relational reasoning abilities than competing contrastive approaches. Finally, we evaluate diffusion models trained on ImageNet and find that they approach the performance of SOTA discriminative classifiers trained on the same dataset, even with weak augmentations and no regularization. Results and visualizations at https://diffusion-classifier.github.io/
翻译:近期大规模文本到图像扩散模型的浪潮极大地提升了我们基于文本的图像生成能力。这些模型能够针对海量提示生成逼真图像,并展现出令人印象深刻的组合泛化能力。迄今为止,几乎所有应用都聚焦于采样;然而,扩散模型还能提供条件密度估计,这对于图像生成之外的任务同样有用。本文表明,像Stable Diffusion这样的大规模文本到图像扩散模型的密度估计,可以在无需任何额外训练的情况下,用于执行零样本分类。我们的生成式分类方法在多种基准测试上取得了强劲结果,并优于从扩散模型中提取知识的其他替代方法。我们还发现,基于扩散的方法在多模态关系推理能力上强于竞争性的对比学习方法。最后,我们评估了在ImageNet上训练的扩散模型,发现即便在弱增强且无正则化的条件下,其性能也接近使用相同数据集训练的最先进判别式分类器。结果与可视化内容见https://diffusion-classifier.github.io/