The landscape of high-performance image generation models is currently shifting from the inefficient multi-step ones to the efficient few-step counterparts (e.g, Z-Image-Turbo and FLUX.2-klein). However, these models present significant challenges for directly continuous supervised fine-tuning. For example, applying the commonly used fine-tuning technique would compromises their inherent few-step inference capability. To address this, we propose D-OPSD, a novel training paradigm for step-distilled diffusion models that enables on-policy learning during supervised fine-tuning. We first find that the modern diffusion model where the LLM/VLM serves as the encoder can inherit its encoder's in-context capabilities. This enables us to make the training as an on-policy self-distillation process. Specifically, during training, we make the model acts as both the teacher and the student with different contexts, where the student is conditioned only on the text feature, while the teacher is conditioned on the multimodal feature of both the text prompt and the target image. Training minimizes the two predicted distributions over the student's own roll-outs. By optimized on the model's own trajectory and under it's own supervision, D-OPSD enables the model to learn new concept, style, etc. without sacrificing the original few-step capacity.


翻译:高性能图像生成模型的格局目前正从低效的多步模型向高效的少步模型(如Z-Image-Turbo和FLUX.2-klein)转变。然而,这些模型在直接进行连续监督微调时面临显著挑战。例如,应用常用的微调技术会损害其固有的少步推理能力。为解决这一问题,我们提出D-OPSD,一种面向步蒸馏扩散模型的新型训练范式,能够在监督微调过程中实现基于策略的学习。我们首先发现,以LLM/VLM为编码器的现代扩散模型能够继承其编码器的上下文能力。这使我们能够将训练过程构建为基于策略的自蒸馏过程。具体而言,在训练阶段,我们让模型在不同上下文中同时扮演教师和学生角色:学生仅以文本特征为条件,而教师则以文本提示和目标图像的多模态特征为条件。训练过程中,模型通过最小化学生自身轨迹上的两个预测分布进行优化。通过基于模型自身轨迹并在其自身监督下进行优化,D-OPSD使模型能够在不牺牲原始少步能力的前提下学习新概念、风格等。

0
下载
关闭预览

相关内容

综述 | OPSD:大语言模型的在线策略自蒸馏
专知会员服务
10+阅读 · 6月1日
扩散模型中的缓存方法综述:迈向高效的多模态生成
专知会员服务
9+阅读 · 2025年10月23日
预训练扩散模型蒸馏综述
专知会员服务
25+阅读 · 2025年2月17日
高效扩散模型:从原理到实践的全面综述
专知会员服务
41+阅读 · 2024年10月16日
多模态可控扩散模型综述
专知会员服务
39+阅读 · 2024年7月20日
低层视觉中的扩散模型:综述
专知会员服务
22+阅读 · 2024年6月18日
去噪扩散概率模型,46页ppt
专知会员服务
63+阅读 · 2023年1月4日
谷歌EfficientNet缩放模型,PyTorch实现登热榜
机器学习算法与Python学习
11+阅读 · 2019年6月4日
深度学习时代的图模型,清华发文综述图网络
GAN生成式对抗网络
13+阅读 · 2018年12月23日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
VIP会员
最新内容
印度精确打击与指挥架构的断层
专知会员服务
4+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
6+阅读 · 7月20日
深入Project Maven:为何人工智能在战场上依然失灵
锻造未来士兵:外骨骼、基因工程与赛博格
专知会员服务
7+阅读 · 7月19日
《无人机蜂群通信技术研究》50页
专知会员服务
10+阅读 · 7月19日
相关VIP内容
综述 | OPSD:大语言模型的在线策略自蒸馏
专知会员服务
10+阅读 · 6月1日
扩散模型中的缓存方法综述:迈向高效的多模态生成
专知会员服务
9+阅读 · 2025年10月23日
预训练扩散模型蒸馏综述
专知会员服务
25+阅读 · 2025年2月17日
高效扩散模型:从原理到实践的全面综述
专知会员服务
41+阅读 · 2024年10月16日
多模态可控扩散模型综述
专知会员服务
39+阅读 · 2024年7月20日
低层视觉中的扩散模型:综述
专知会员服务
22+阅读 · 2024年6月18日
去噪扩散概率模型,46页ppt
专知会员服务
63+阅读 · 2023年1月4日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员