Image editing using diffusion models has witnessed extremely fast-paced growth recently. There are various ways in which previous works enable controlling and editing images. Some works use high-level conditioning such as text, while others use low-level conditioning. Nevertheless, most of them lack fine-grained control over the properties of the different objects present in the image, i.e. object-level image editing. In this work, we consider an image as a composition of multiple objects, each defined by various properties. Out of these properties, we identify structure and appearance as the most intuitive to understand and useful for editing purposes. We propose Structure-and-Appearance Paired Diffusion model (PAIR-Diffusion), which is trained using structure and appearance information explicitly extracted from the images. The proposed model enables users to inject a reference image's appearance into the input image at both the object and global levels. Additionally, PAIR-Diffusion allows editing the structure while maintaining the style of individual components of the image unchanged. We extensively evaluate our method on LSUN datasets and the CelebA-HQ face dataset, and we demonstrate fine-grained control over both structure and appearance at the object level. We also applied the method to Stable Diffusion to edit any real image at the object level.
翻译:近年来,利用扩散模型进行图像编辑的技术发展极为迅速。以往研究提供了多种控制与编辑图像的方式,部分工作采用文本等高层条件约束,另一部分则利用低层条件。然而,大多数方法缺乏对图像中不同物体属性的细粒度控制,即无法实现物体级图像编辑。本文将图像视为多个物体的组合,每个物体由不同属性定义。在这些属性中,我们识别出结构与外观是最直观且最有利于编辑的属性。我们提出结构与外观配对扩散模型(PAIR-Diffusion),该模型利用从图像中显式提取的结构与外观信息进行训练。所提模型支持用户将参考图像的外观注入输入图像中,既能实现物体级注入,也能实现全局级注入。此外,PAIR-Diffusion允许在保持图像中各组件风格不变的前提下编辑结构。我们在LSUN数据集和CelebA-HQ人脸数据集上对本方法进行了广泛评估,展示了在物体级对结构与外观的细粒度控制能力。我们还将该方法应用于Stable Diffusion,实现了对任意真实图像的物体级编辑。