Blind super-resolution methods based on stable diffusion showcase formidable generative capabilities in reconstructing clear high-resolution images with intricate details from low-resolution inputs. However, their practical applicability is often hampered by poor efficiency, stemming from the requirement of thousands or hundreds of sampling steps. Inspired by the efficient text-to-image approach adversarial diffusion distillation (ADD), we design AddSR to address this issue by incorporating the ideas of both distillation and ControlNet. Specifically, we first propose a prediction-based self-refinement strategy to provide high-frequency information in the student model output with marginal additional time cost. Furthermore, we refine the training process by employing HR images, rather than LR images, to regulate the teacher model, providing a more robust constraint for distillation. Second, we introduce a timestep-adapting loss to address the perception-distortion imbalance problem introduced by ADD. Extensive experiments demonstrate our AddSR generates better restoration results, while achieving faster speed than previous SD-based state-of-the-art models (e.g., 7x faster than SeeSR).
翻译:基于稳定扩散的盲超分辨率方法在从低分辨率输入重建包含复杂细节的清晰高分辨率图像方面展现出强大的生成能力。然而,其实际应用常因需要数千或数百个采样步骤而受限于低效性。受高效文本到图像方法对抗扩散蒸馏(ADD)启发,我们设计AddSR,融合蒸馏与ControlNet思想以解决该问题。具体而言,我们首先提出一种基于预测的自精炼策略,以极小的额外时间成本在学生模型输出中提供高频信息。其次,我们优化训练过程,采用高分辨率图像而非低分辨率图像来约束教师模型,从而为蒸馏提供更稳健的约束。第三,我们引入时序自适应损失以解决ADD带来的感知-失真不平衡问题。大量实验表明,AddSR在生成更优复原结果的同时,相较此前基于稳定扩散的最先进模型实现了更快速度(例如比SeeSR快7倍)。