Diffusion models are typically trained with objectives that focus on local denoising targets at individual time steps (or adjacent pairs), which do not enforce consistency between predictions along the denoising trajectory. This lack of cross-time consistency can degrade performance, especially for few-step samplers. We introduce a temporal difference (TD) objective that penalizes inconsistency of the model's multi-step progress along the denoising path. By reformulating the diffusion process as a Markov reward process and casting denoising as a policy evaluation problem in reinforcement learning, we derive a unified TD approach that applies to both discrete- and continuous-time diffusion formulations. We further propose a principled sample-based reweighting method that stabilizes training. Empirically, we show that using our TD training can significantly improve sample quality measured by FID, with stronger advantages when the number of sampling steps is small, highlighting its practical utility under low-computation-budget scenarios. We provide ablation studies to justify our design choices, including pairwise loss reweighting, regularization weight, and one-step stride. Overall, our TD approach can be a general drop-in that enforces cross-time consistency and improves generation quality across different diffusion generative models.


翻译:扩散模型通常通过聚焦于单个时间步(或相邻步)的局部去噪目标进行训练,这种做法并未强制要求在去噪轨迹上预测的一致性。这种跨时间一致性的缺失会降低模型性能,尤其是在少步采样器中。我们提出一种时间差分(TD)目标函数,通过惩罚模型沿去噪路径的多步进展不一致性来解决此问题。通过将扩散过程重新表述为马尔可夫奖励过程,并将去噪视为强化学习中的策略评估问题,我们推导出一种统一的TD方法,可同时适用于离散时间和连续时间扩散公式。此外,我们提出一种基于样本的原则性重加权方法以稳定训练。实验表明,使用我们的TD训练能显著提升由FID衡量的样本质量,且在采样步数较少时优势更为突出,凸显其在低计算预算场景下的实用价值。我们通过消融研究验证了设计选择的合理性,包括成对损失重加权、正则化权重及单步跨度。总体而言,我们的TD方法可作为一种通用即插即用模块,通过强制跨时间一致性来提升各类扩散生成模型的生成质量。

0
下载
关闭预览

相关内容

用于强化学习的扩散模型:基础、分类与发展
专知会员服务
24+阅读 · 2025年10月15日
用于时间序列预测的扩散模型:综述
专知会员服务
30+阅读 · 2025年7月22日
时间序列和时空数据扩散模型综述
专知会员服务
64+阅读 · 2024年5月1日
视觉的有效扩散模型综述
专知会员服务
97+阅读 · 2022年10月20日
深度学习中Attention Mechanism详细介绍:原理、分类及应用
深度学习与NLP
10+阅读 · 2019年2月18日
迁移学习在深度学习中的应用
专知
24+阅读 · 2017年12月24日
图上的归纳表示学习
科技创新与创业
23+阅读 · 2017年11月9日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
16+阅读 · 2013年12月31日
VIP会员
最新内容
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 今天4:08
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
“史诗怒火”行动:现代多域作战的重要节点
专知会员服务
8+阅读 · 7月30日
《下一代无线网络中的多无人机通信资源管理》
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
16+阅读 · 2013年12月31日
Top
微信扫码咨询专知VIP会员