The ability to fine-tune generative models for text-to-image generation tasks is crucial, particularly facing the complexity involved in accurately interpreting and visualizing textual inputs. While LoRA is efficient for language model adaptation, it often falls short in text-to-image tasks due to the intricate demands of image generation, such as accommodating a broad spectrum of styles and nuances. To bridge this gap, we introduce StyleInject, a specialized fine-tuning approach tailored for text-to-image models. StyleInject comprises multiple parallel low-rank parameter matrices, maintaining the diversity of visual features. It dynamically adapts to varying styles by adjusting the variance of visual features based on the characteristics of the input signal. This approach significantly minimizes the impact on the original model's text-image alignment capabilities while adeptly adapting to various styles in transfer learning. StyleInject proves particularly effective in learning from and enhancing a range of advanced, community-fine-tuned generative models. Our comprehensive experiments, including both small-sample and large-scale data fine-tuning as well as base model distillation, show that StyleInject surpasses traditional LoRA in both text-image semantic consistency and human preference evaluation, all while ensuring greater parameter efficiency.
翻译:摘要:针对文本到图像生成任务中精确理解与可视化文本输入的复杂性,微调生成模型的能力至关重要。尽管LoRA在语言模型适配中表现高效,但由于图像生成需处理广泛风格与细微差异等复杂需求,其在文本到图像任务中往往力有不逮。为此,我们提出专为文本到图像模型设计的微调方法StyleInject。该方法包含多个并行低秩参数矩阵以保持视觉特征多样性,通过依据输入信号特征调整视觉特征方差实现风格动态适配。该策略在保持原始模型文本-图像对齐能力的同时,能灵活适应迁移学习中的各类风格。StyleInject在基于社区微调的高级生成模型学习与增强中尤为有效。涵盖小样本与大规模数据微调及基座模型蒸馏的综合实验表明:StyleInject在文本-图像语义一致性与人类偏好评估中均超越传统LoRA,且参数效率更优。