Diffusion models achieve state-of-the-art image synthesis, with their generative trajectories fundamentally exhibiting a spectral bias, resolving low-frequency global structures early and high-frequency fine details later. Conventional stochastic differential equation (SDE) solvers fail to account for this dynamic, naively injecting uniform white noise throughout the entire process and misusing the finite energy budget. In this work, we establish a mathematical framework that reconsiders SDE inference as a targeted, frequency-decoupled energy transfer. Leveraging this framework, we introduce Colored Noise Sampling (CNS), a novel, training-free stochastic solver. Rather than injecting uniform white noise, CNS utilizes a dynamic, timestep- and frequency-dependent schedule that more efficiently allocates injected energy toward structurally unresolved frequency bands. By actively exploiting the model's inherent spectral bias, CNS systematically steers the generated distribution toward the true data manifold. Extensive experiments demonstrate that CNS significantly outperforms standard ODE and SDE baselines as a strictly plug-and-play, inference-time sampler substitution across diverse architectures (SiT, JiT, FLUX). Compared to standard sampling on ImageNet-256, CNS achieves substantial unguided FID reductions, improving from 8.26 to 6.27 on SiT-XL/2, 32.39 to 26.69 on JiT-B/16, and 11.88 to 8.31 on JiT-H/16, while yielding consistent relative FID improvements with Classifier-Free Guidance. Project page is available at https://hadardavidson.github.io/CNS/.


翻译:扩散模型在图像合成中实现了最先进的性能,其生成轨迹本质上表现出谱偏差,即早期解析低频全局结构,后期处理高频细节。传统的随机微分方程求解器未能考虑这一动态特性,在整个过程中简单地注入均匀白噪声,并错误地使用有限能量预算。本研究建立了一个数学框架,将随机微分方程推理重新视为一种针对性的、频率解耦的能量传递。利用此框架,我们提出了彩色噪声采样(CNS),一种新颖的、无需训练随机求解器。不同于注入均匀白噪声,CNS采用动态的、与时间步和频率相关的调度策略,更高效地将注入能量分配给结构上未解析的频率波段。通过主动利用模型固有的谱偏差,CNS系统地引导生成分布趋向真实数据流形。大量实验表明,作为严格的即插即用推理阶段采样器替代方案,CNS在多种架构(SiT、JiT、FLUX)上显著优于标准常微分方程和随机微分方程基线。在ImageNet-256标准采样中,无需引导的FID分数大幅降低:SiT-XL/2从8.26提升至6.27,JiT-B/16从32.39提升至26.69,JiT-H/16从11.88提升至8.31,且在无分类器引导下取得一致的相对FID改进。项目页面:https://hadardavidson.github.io/CNS/。

0
下载
关闭预览

相关内容

144页ppt《扩散模型》,Google DeepMind Sander Dieleman
专知会员服务
51+阅读 · 2025年11月21日
用于语言生成的离散扩散模型
专知会员服务
12+阅读 · 2025年7月10日
《扩散模型图像编辑》综述
专知会员服务
28+阅读 · 2024年2月28日
扩散模型图像超分辨率等综述
专知会员服务
25+阅读 · 2024年1月2日
去噪扩散概率模型,46页ppt
专知会员服务
63+阅读 · 2023年1月4日
详解扩散模型:从DDPM到稳定扩散,附Slides与视频
专知会员服务
87+阅读 · 2022年10月9日
CVPR 2019 | 无监督领域特定单图像去模糊
PaperWeekly
14+阅读 · 2019年3月20日
CVPR 2018 论文解读 | 基于GAN和CNN的图像盲去噪
PaperWeekly
13+阅读 · 2019年1月22日
Image Captioning 36页最新综述, 161篇参考文献
专知
90+阅读 · 2018年10月23日
图像降噪算法介绍及实现汇总
极市平台
26+阅读 · 2018年1月3日
最新|深度离散哈希算法,可用于图像检索!
全球人工智能
14+阅读 · 2017年12月15日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Arxiv
0+阅读 · 6月12日
Arxiv
0+阅读 · 6月9日
Arxiv
0+阅读 · 5月6日
VIP会员
最新内容
边缘计算的军事应用
专知会员服务
7+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
9+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
14+阅读 · 8月7日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员