While diffusion distillation has enabled one-step generation through methods like Variational Score Distillation, adapting distilled models to emerging new controls -- such as novel structural constraints or latest user preferences -- remains challenging. Conventional approaches typically requires modifying the base diffusion model and redistilling it -- a process that is both computationally intensive and time-consuming. To address these challenges, we introduce Joint Distribution Matching (JDM), a novel approach that minimizes the reverse KL divergence between image-condition joint distributions. By deriving a tractable upper bound, JDM decouples fidelity learning from condition learning. This asymmetric distillation scheme enables our one-step student to handle controls unknown to the teacher model and facilitates improved classifier-free guidance (CFG) usage and seamless integration of human feedback learning (HFL). Experimental results demonstrate that JDM surpasses baseline methods such as multi-step ControlNet by mere one-step in most cases, while achieving state-of-the-art performance in one-step text-to-image synthesis through improved usage of CFG or HFL integration.
翻译:尽管扩散蒸馏技术已通过变分分数蒸馏等方法实现了单步生成,但使蒸馏模型适应新兴控制条件——例如新型结构约束或最新用户偏好——仍然具有挑战性。传统方法通常需要修改基础扩散模型并重新进行蒸馏,这一过程计算密集且耗时。为应对这些挑战,我们提出了联合分布匹配方法,该方法通过最小化图像-条件联合分布的反向KL散度来实现控制适配。通过推导可处理的上界,JDM将保真度学习与条件学习解耦。这种非对称蒸馏方案使我们的单步学生模型能够处理教师模型未知的控制条件,并促进改进的无分类器引导使用及人类反馈学习的无缝集成。实验结果表明,在多数情况下JDM仅需单步推理即可超越多步ControlNet等基线方法,同时通过改进的CFG使用或HFL集成,在单步文本到图像合成任务中达到了最先进的性能。