Denoising diffusion models enable conditional generation and density modeling of complex relationships like images and text. However, the nature of the learned relationships is opaque making it difficult to understand precisely what relationships between words and parts of an image are captured, or to predict the effect of an intervention. We illuminate the fine-grained relationships learned by diffusion models by noticing a precise relationship between diffusion and information decomposition. Exact expressions for mutual information and conditional mutual information can be written in terms of the denoising model. Furthermore, pointwise estimates can be easily estimated as well, allowing us to ask questions about the relationships between specific images and captions. Decomposing information even further to understand which variables in a high-dimensional space carry information is a long-standing problem. For diffusion models, we show that a natural non-negative decomposition of mutual information emerges, allowing us to quantify informative relationships between words and pixels in an image. We exploit these new relations to measure the compositional understanding of diffusion models, to do unsupervised localization of objects in images, and to measure effects when selectively editing images through prompt interventions.
翻译:去噪扩散模型能够实现条件生成以及对图像与文本等复杂关系的密度建模。然而,所学关系的本质是不透明的,这导致难以精确理解图像中词汇与像素区域之间到底捕获了何种关系,也难以预测干预操作的影响。我们通过揭示扩散与信息分解之间的精确关系,阐明了扩散模型所学习的细粒度关系。互信息和条件互信息可以用去噪模型写成精确表达式。此外,逐点估计也可以轻松计算,从而可以探究特定图像与描述之间的关系。进一步将信息分解以理解高维空间中哪些变量携带信息,是一个长期存在的难题。对于扩散模型,我们证明了一种自然的非负互信息分解方法得以涌现,使我们能够量化图像中词汇与像素之间的信息关系。我们利用这些新关系来评估扩散模型的组合理解能力,实现图像中对象的无监督定位,并通过提示干预测量选择性编辑图像时的影响。