Diffusion models have shown significant progress in image translation tasks recently. However, due to their stochastic nature, there's often a trade-off between style transformation and content preservation. Current strategies aim to disentangle style and content, preserving the source image's structure while successfully transitioning from a source to a target domain under text or one-shot image conditions. Yet, these methods often require computationally intense fine-tuning of diffusion models or additional neural networks. To address these challenges, here we present an approach that guides the reverse process of diffusion sampling by applying asymmetric gradient guidance. This results in quicker and more stable image manipulation for both text-guided and image-guided image translation. Our model's adaptability allows it to be implemented with both image- and latent-diffusion models. Experiments show that our method outperforms various state-of-the-art models in image translation tasks.
翻译:扩散模型在图像翻译任务中近年来取得了显著进展。然而,由于其随机性,风格转换与内容保留之间往往存在权衡。当前策略旨在解耦风格与内容,在文本或单张图像条件下,保留源图像结构的同时成功实现从源域到目标域的过渡。然而,这些方法通常需要对扩散模型进行计算密集型的微调,或额外增加神经网络。为解决这些挑战,本文提出了一种方法,通过应用非对称梯度引导来指导扩散采样的逆过程。这实现了在文本引导和图像引导的图像翻译中更快速、更稳定的图像操作。我们模型的适应性使其能够同时应用于图像扩散模型和潜在扩散模型。实验表明,我们的方法在图像翻译任务中优于多种现有最先进模型。