The substantial computational costs of diffusion models, especially due to the repeated denoising steps necessary for high-quality image generation, present a major obstacle to their widespread adoption. While several studies have attempted to address this issue by reducing the number of score function evaluations (NFE) using advanced ODE solvers without fine-tuning, the decreased number of denoising iterations misses the opportunity to update fine details, resulting in noticeable quality degradation. In our work, we introduce an advanced acceleration technique that leverages the temporal redundancy inherent in diffusion models. Reusing feature maps with high temporal similarity opens up a new opportunity to save computation resources without compromising output quality. To realize the practical benefits of this intuition, we conduct an extensive analysis and propose a novel method, FRDiff. FRDiff is designed to harness the advantages of both reduced NFE and feature reuse, achieving a Pareto frontier that balances fidelity and latency trade-offs in various generative tasks.
翻译:扩散模型高昂的计算成本,特别是高质量图像生成所需的重复去噪步骤,成为其广泛部署的主要障碍。尽管多项研究尝试通过使用先进ODE求解器在不微调情况下减少评分函数评估次数(NFE)来缓解此问题,但去噪迭代次数的减少会遗漏细节更新的机会,导致可见的质量退化。本文提出一种利用扩散模型固有时间冗余性的先进加速技术。复用具有高度时间相似性的特征图为节省计算资源而不损害输出质量提供了新契机。为实现该直觉的实际效益,我们进行了广泛分析并提出新型方法FRDiff。FRDiff旨在融合减少NFE与特征复用的双重优势,在各类生成任务中实现了平衡保真度与延迟的帕累托前沿。