Text-conditional diffusion models generate high-quality, diverse images. However, text is often an ambiguous specification for a desired target image, creating the need for additional user-friendly controls for diffusion-based image generation. We focus on having precise control over image output for scenes with several objects. Users control image generation by defining a collage: a text prompt paired with an ordered sequence of layers, where each layer is an RGBA image and a corresponding text prompt. We introduce Collage Diffusion, a collage-conditional diffusion algorithm that allows users to control both the spatial arrangement and visual attributes of objects in the scene, and also enables users to edit individual components of generated images. To ensure that different parts of the input text correspond to the various locations specified in the input collage layers, Collage Diffusion modifies text-image cross-attention with the layers' alpha masks. To maintain characteristics of individual collage layers that are not specified in text, Collage Diffusion learns specialized text representations per layer. Collage input also enables layer-based controls that provide fine-grained control over the final output: users can control image harmonization on a layer-by-layer basis, and they can edit individual objects in generated images while keeping other objects fixed. Collage-conditional image generation requires harmonizing the input collage to make objects fit together--the key challenge involves minimizing changes in the positions and key visual attributes of objects in the input collage while allowing other attributes of the collage to change in the harmonization process. By leveraging the rich information present in layer input, Collage Diffusion generates globally harmonized images that maintain desired object locations and visual characteristics better than prior approaches.
翻译:文本条件扩散模型能够生成高质量、多样化的图像。然而,文本对目标图像的描述往往存在歧义,这促使我们需要为基于扩散的图像生成提供更多用户友好的控制方式。本文聚焦于对包含多个物体的场景实现精确的图像输出控制。用户通过定义拼贴来指导图像生成:即一个文本提示与有序图层序列的组合,其中每个图层包含一张RGBA图像及其对应的文本提示。我们提出拼贴扩散(Collage Diffusion),一种基于拼贴条件的扩散算法,该算法允许用户控制场景中物体的空间布局与视觉属性,并可编辑生成图像中的独立组件。为确保输入文本的不同部分与拼贴图层中指定的位置相对应,拼贴扩散利用图层的Alpha掩码修正文本-图像交叉注意力机制。为保留文本未指定的单个拼贴图层特征,拼贴扩散为每个图层学习专用文本表征。拼贴输入还支持基于图层的控制,使用户能精细调控最终输出:用户可逐层控制图像融合程度,并在保持其他物体不变的前提下编辑生成图像中的独立物体。基于拼贴条件的图像生成需要融合输入拼贴以协调各物体——核心挑战在于最小化输入拼贴中物体位置与关键视觉属性的变化,同时允许拼贴的其他属性在协调过程中发生改变。通过充分利用图层输入中的丰富信息,拼贴扩散生成的全局协调图像在保持期望物体位置与视觉特征方面优于现有方法。