Text-to-image models (T2I) offer a new level of flexibility by allowing users to guide the creative process through natural language. However, personalizing these models to align with user-provided visual concepts remains a challenging problem. The task of T2I personalization poses multiple hard challenges, such as maintaining high visual fidelity while allowing creative control, combining multiple personalized concepts in a single image, and keeping a small model size. We present Perfusion, a T2I personalization method that addresses these challenges using dynamic rank-1 updates to the underlying T2I model. Perfusion avoids overfitting by introducing a new mechanism that "locks" new concepts' cross-attention Keys to their superordinate category. Additionally, we develop a gated rank-1 approach that enables us to control the influence of a learned concept during inference time and to combine multiple concepts. This allows runtime-efficient balancing of visual-fidelity and textual-alignment with a single 100KB trained model, which is five orders of magnitude smaller than the current state of the art. Moreover, it can span different operating points across the Pareto front without additional training. Finally, we show that Perfusion outperforms strong baselines in both qualitative and quantitative terms. Importantly, key-locking leads to novel results compared to traditional approaches, allowing to portray personalized object interactions in unprecedented ways, even in one-shot settings.
翻译:摘要:文本到图像模型(T2I)通过允许用户通过自然语言引导创作过程,提供了前所未有的灵活性。然而,将这些模型个性化以适应用户提供的视觉概念仍是一项具有挑战性的问题。T2I个性化任务面临多重困难,例如在保持高视觉保真度的同时实现创意控制、在单张图像中融合多个个性化概念,以及维持较小的模型规模。本文提出Perfusion——一种通过动态秩一更新对底层T2I模型进行个性化的方法。该方法引入一种新机制,将新概念的交叉注意力键“锁定”到其上级类别,从而避免过拟合。此外,我们开发了一种门控秩一方法,可在推理阶段控制所学概念的影响,并融合多个概念。这使得通过单个100KB的训练模型(比当前最优模型小五个数量级)实现运行时高效的视觉保真度与文本对齐平衡,且无需额外训练即可在帕累托前沿覆盖不同操作点。最后,我们证明Perfusion在定性与定量评估中均优于强基线方法。重要的是,与经典方法相比,键锁定机制带来了新颖的成果,能够在单样本场景中以前所未有的方式呈现个性化对象的交互。