The incredible generative ability of large-scale text-to-image (T2I) models has demonstrated strong power of learning complex structures and meaningful semantics. However, relying solely on text prompts cannot fully take advantage of the knowledge learned by the model, especially when flexible and accurate controlling (e.g., color and structure) is needed. In this paper, we aim to ``dig out" the capabilities that T2I models have implicitly learned, and then explicitly use them to control the generation more granularly. Specifically, we propose to learn simple and lightweight T2I-Adapters to align internal knowledge in T2I models with external control signals, while freezing the original large T2I models. In this way, we can train various adapters according to different conditions, achieving rich control and editing effects in the color and structure of the generation results. Further, the proposed T2I-Adapters have attractive properties of practical value, such as composability and generalization ability. Extensive experiments demonstrate that our T2I-Adapter has promising generation quality and a wide range of applications.
翻译:大规模文本到图像(T2I)模型展现出了学习复杂结构与有意义语义的强大生成能力。然而,仅依赖文本提示无法充分利用模型所习得的知识,尤其是在需要灵活且精确的控制(如颜色与结构)时。本文旨在"挖掘"T2I模型已隐式学习的能力,并将其显式用于更细粒度的生成控制。具体而言,我们提出学习简单轻量的T2I-Adapter,在冻结原始大规模T2I模型的同时,将其内部知识与外部控制信号对齐。通过这种方式,我们能够根据不同的条件训练多种适配器,实现对生成结果在颜色与结构上的丰富控制与编辑效果。此外,所提出的T2I-Adapter具有实用的组合性与泛化能力等特性。大量实验表明,我们的T2I-Adapter具有出色的生成质量与广泛的应用场景。