Training-free guided sampling in diffusion models leverages off-the-shelf pre-trained networks, such as an aesthetic evaluation model, to guide the generation process. Current training-free guided sampling algorithms obtain the guidance energy function based on a one-step estimate of the clean image. However, since the off-the-shelf pre-trained networks are trained on clean images, the one-step estimation procedure of the clean image may be inaccurate, especially in the early stages of the generation process in diffusion models. This causes the guidance in the early time steps to be inaccurate. To overcome this problem, we propose Symplectic Adjoint Guidance (SAG), which calculates the gradient guidance in two inner stages. Firstly, SAG estimates the clean image via $n$ function calls, where $n$ serves as a flexible hyperparameter that can be tailored to meet specific image quality requirements. Secondly, SAG uses the symplectic adjoint method to obtain the gradients accurately and efficiently in terms of the memory requirements. Extensive experiments demonstrate that SAG generates images with higher qualities compared to the baselines in both guided image and video generation tasks.
翻译:无训练引导采样在扩散模型中利用现成预训练网络(如美学评估模型)来指导生成过程。现有无训练引导采样算法基于一步估计的干净图像获取引导能量函数。然而,由于现成预训练网络在干净图像上训练,一步估计干净图像的过程可能不准确,尤其在扩散模型生成过程的早期阶段,导致初始时间步的引导出现偏差。为解决这一问题,我们提出辛伴随引导(SAG),该方法通过两个内阶段计算梯度引导。首先,SAG通过$n$次函数调用估计干净图像,其中$n$作为灵活超参数,可根据特定图像质量需求进行调整。其次,SAG利用辛伴随方法在内存需求方面高效且精确地获取梯度。大量实验表明,在引导图像和视频生成任务中,SAG生成的图像质量均优于基线方法。