Diffusion models have emerged as powerful tools for high-quality image generation and editing, but guiding these models to produce specific outputs remains a challenge. Conventional approaches rely on conditioning mechanisms, such as text prompts or semantic maps, which require extensively annotated datasets. In this preliminary work, we explore diffusion models conditioned on representations from a pre-trained self-supervised model. The self-conditioning mechanism not only improves the quality of unconditional image generation, but also provides a representation space that can be used to control the generation. We explore this conditioning space by identifying directions of variations, and demonstrate promising properties in terms of smoothness and disentanglement.
翻译:扩散模型已成为高质量图像生成与编辑的强大工具,但引导这些模型产生特定输出仍具挑战性。传统方法依赖于文本提示或语义地图等条件机制,需要大量标注数据集。在本初步研究中,我们探索了基于预训练自监督模型表示进行条件控制的扩散模型。该自条件机制不仅提升了无条件图像生成的质量,还提供了可用于控制生成过程的表示空间。我们通过识别变分方向来探索该条件空间,并展示了其在平滑性和解耦性方面的优良特性。