Blind super-resolution methods based on stable diffusion showcase formidable generative capabilities in reconstructing clear high-resolution images with intricate details from low-resolution inputs. However, their practical applicability is often hampered by poor efficiency, stemming from the requirement of thousands or hundreds of sampling steps. Inspired by the efficient adversarial diffusion distillation (ADD), we design~\name~to address this issue by incorporating the ideas of both distillation and ControlNet. Specifically, we first propose a prediction-based self-refinement strategy to provide high-frequency information in the student model output with marginal additional time cost. Furthermore, we refine the training process by employing HR images, rather than LR images, to regulate the teacher model, providing a more robust constraint for distillation. Second, we introduce a timestep-adaptive ADD to address the perception-distortion imbalance problem introduced by original ADD. Extensive experiments demonstrate our~\name~generates better restoration results, while achieving faster speed than previous SD-based state-of-the-art models (e.g., $7$$\times$ faster than SeeSR).
翻译:基于稳定扩散的盲超分辨率方法在从低分辨率输入重建具有复杂细节的清晰高分辨率图像方面展现出强大的生成能力。然而,其实际应用常受限于较差的效率,这源于需要数千或数百次采样步骤。受高效的对抗扩散蒸馏(ADD)启发,我们设计了AddSR,通过结合蒸馏和ControlNet的思想来解决这一问题。具体而言,我们首先提出一种基于预测的自优化策略,以边际额外时间成本为学生模型输出提供高频信息。此外,我们通过使用高分辨率图像而非低分辨率图像来调控教师模型,从而优化训练过程,为蒸馏提供更稳健的约束。其次,我们引入时间步自适应ADD,以解决原始ADD带来的感知-失真不平衡问题。大量实验表明,我们的AddSR能生成更好的复原结果,同时比以往基于SD的最先进模型实现更快的速度(例如,比SeeSR快7倍)。