Personalizing generative models offers a way to guide image generation with user-provided references. Current personalization methods can invert an object or concept into the textual conditioning space and compose new natural sentences for text-to-image diffusion models. However, representing and editing specific visual attributes like material, style, layout, etc. remains a challenge, leading to a lack of disentanglement and editability. To address this, we propose a novel approach that leverages the step-by-step generation process of diffusion models, which generate images from low- to high-frequency information, providing a new perspective on representing, generating, and editing images. We develop Prompt Spectrum Space P*, an expanded textual conditioning space, and a new image representation method called ProSpect. ProSpect represents an image as a collection of inverted textual token embeddings encoded from per-stage prompts, where each prompt corresponds to a specific generation stage (i.e., a group of consecutive steps) of the diffusion model. Experimental results demonstrate that P* and ProSpect offer stronger disentanglement and controllability compared to existing methods. We apply ProSpect in various personalized attribute-aware image generation applications, such as image/text-guided material/style/layout transfer/editing, achieving previously unattainable results with a single image input without fine-tuning the diffusion models.
翻译:个性化生成模型为用户提供的参考引导图像生成提供了途径。现有个性化方法可将对象或概念反转到文本条件空间,并通过组合自然语言句子驱动文本到图像扩散模型。然而,对材质、风格、布局等特定视觉属性的表征与编辑仍是难题,导致解耦性与可编辑性不足。为此,我们提出了一种新方法,利用扩散模型从低频到高频信息逐步生成的特性,为图像表征、生成与编辑提供了新视角。具体而言,我们构建了扩展文本条件空间——提示谱空间P*,并提出名为ProSpect的新图像表征方法。ProSpect将图像表征为从逐阶段提示中编码的反转文本词嵌入集合,其中每个提示对应扩散模型的一个特定生成阶段(即连续步骤组)。实验结果表明,与现有方法相比,P*与ProSpect具有更强的解耦性与可控性。我们已将ProSpect应用于多种个性化属性感知图像生成场景,例如基于图像/文本引导的材质/风格/布局迁移与编辑,在无需微调扩散模型且仅需单张图像输入的情况下,实现了此前不可达的结果。