Diffusion models have shown superior performance in image generation and manipulation, but the inherent stochasticity presents challenges in preserving and manipulating image content and identity. While previous approaches like DreamBooth and Textual Inversion have proposed model or latent representation personalization to maintain the content, their reliance on multiple reference images and complex training limits their practicality. In this paper, we present a simple yet highly effective approach to personalization using highly personalized (HiPer) text embedding by decomposing the CLIP embedding space for personalization and content manipulation. Our method does not require model fine-tuning or identifiers, yet still enables manipulation of background, texture, and motion with just a single image and target text. Through experiments on diverse target texts, we demonstrate that our approach produces highly personalized and complex semantic image edits across a wide range of tasks. We believe that the novel understanding of the text embedding space presented in this work has the potential to inspire further research across various tasks.
翻译:扩散模型在图像生成与编辑领域展现出卓越性能,但其固有的随机性给保持和操控图像内容及身份特征带来挑战。尽管DreamBooth和文本反演等先前方法通过模型或隐空间表示个性化来维持内容一致性,但它们对多张参考图像和复杂训练的依赖限制了实际应用价值。本文提出一种简洁高效的个性化方法——高度个性化(HiPer)文本嵌入,通过分解CLIP嵌入空间实现个性化与内容操控。该方法无需模型微调或标识符,仅凭单张图像与目标文本即可实现背景、纹理与运动等属性的编辑。针对多样化目标文本的实验表明,本方法能在广泛任务中生成高度个性化且语义复杂的图像编辑。我们相信,本文对文本嵌入空间的新颖理解将激发相关领域后续研究。