We present a novel method for exemplar-based image translation, called matching interleaved diffusion models (MIDMs). Most existing methods for this task were formulated as GAN-based matching-then-generation framework. However, in this framework, matching errors induced by the difficulty of semantic matching across cross-domain, e.g., sketch and photo, can be easily propagated to the generation step, which in turn leads to degenerated results. Motivated by the recent success of diffusion models overcoming the shortcomings of GANs, we incorporate the diffusion models to overcome these limitations. Specifically, we formulate a diffusion-based matching-and-generation framework that interleaves cross-domain matching and diffusion steps in the latent space by iteratively feeding the intermediate warp into the noising process and denoising it to generate a translated image. In addition, to improve the reliability of the diffusion process, we design a confidence-aware process using cycle-consistency to consider only confident regions during translation. Experimental results show that our MIDMs generate more plausible images than state-of-the-art methods.
翻译:我们提出了一种新颖的示例驱动图像翻译方法,称为匹配交错扩散模型(MIDMs)。现有的大多数方法基于生成对抗网络实现了"匹配-生成"框架,但该框架中跨域语义匹配(如素描与照片)的困难会导致匹配误差易传播至生成步骤,进而引发退化结果。受扩散模型近年来在克服生成对抗网络局限性方面取得的成功启发,我们将扩散模型引入以突破这些限制。具体而言,我们构建了基于扩散的"匹配-生成"框架,通过在潜在空间中将中间扭曲结果迭代注入噪声过程并进行去噪生成翻译图像,实现跨域匹配与扩散步骤的交错处理。此外,为提升扩散过程的可靠性,我们基于循环一致性设计了置信度感知流程,仅在置信区域执行翻译。实验结果表明,与现有最优方法相比,我们的MIDMs能生成更逼真的图像。