Stylized text-to-image generation focuses on creating images from textual descriptions while adhering to a style specified by a few reference images. However, subtle style variations within different reference images can hinder the model from accurately learning the target style. In this paper, we propose InstaStyle, a novel approach that excels in generating high-fidelity stylized images with only a single reference image. Our approach is based on the finding that the inversion noise from a stylized reference image inherently carries the style signal, as evidenced by their non-zero signal-to-noise ratio. We employ DDIM inversion to extract this noise from the reference image and leverage a diffusion model to generate new stylized images from the "style" noise. Additionally, the inherent ambiguity and bias of textual prompts impede the precise conveying of style. To address this, we introduce a learnable style token via prompt refinement, which enhances the accuracy of the style description for the reference image. Qualitative and quantitative experimental results demonstrate that InstaStyle achieves superior performance compared to current benchmarks. Furthermore, our approach also showcases its capability in the creative task of style combination with mixed inversion noise.
翻译:风格化文本到图像生成旨在根据文本描述,同时遵循由少量参考图像指定的风格来生成图像。然而,不同参考图像中微妙的风格差异会妨碍模型精确学习目标风格。本文提出InstaStyle,一种仅需单张参考图像即可生成高保真风格化图像的新颖方法。我们的方法基于以下发现:风格化参考图像中的逆噪声本质上承载着风格信号,其非零信噪比证实了这一点。我们采用DDIM逆变换从参考图像中提取该噪声,并利用扩散模型从“风格”噪声生成新的风格化图像。此外,文本提示固有的模糊性和偏差阻碍了风格的精确传达。为解决此问题,我们通过提示优化引入可学习的风格标记,从而增强对参考图像风格描述的准确性。定性和定量实验结果表明,InstaStyle相较于现有基准方法取得了更优性能。同时,我们的方法还展示了其在混合逆噪声进行风格组合这一创造性任务中的能力。