Diffusion models have demonstrated exceptional efficacy in various generative applications. While existing models focus on minimizing a weighted sum of denoising score matching losses for data distribution modeling, their training primarily emphasizes instance-level optimization, overlooking valuable structural information within each mini-batch, indicative of pair-wise relationships among samples. To address this limitation, we introduce Structure-guided Adversarial training of Diffusion Models (SADM). In this pioneering approach, we compel the model to learn manifold structures between samples in each training batch. To ensure the model captures authentic manifold structures in the data distribution, we advocate adversarial training of the diffusion generator against a novel structure discriminator in a minimax game, distinguishing real manifold structures from the generated ones. SADM substantially improves existing diffusion transformers (DiT) and outperforms existing methods in image generation and cross-domain fine-tuning tasks across 12 datasets, establishing a new state-of-the-art FID of 1.58 and 2.11 on ImageNet for class-conditional image generation at resolutions of 256x256 and 512x512, respectively.
翻译:扩散模型在各种生成性应用中都展现出卓越的效能。尽管现有模型侧重于通过最小化加权去噪评分匹配损失来对数据分布建模,其训练主要强调实例层面的优化,忽视了每个小批量中隐含的、表征样本间成对关系的宝贵结构信息。为解决这一局限,我们提出了扩散模型的结构引导对抗训练(SADM)。在这一开创性方法中,我们迫使模型学习每个训练批次中样本间的流形结构。为确保模型捕捉到数据分布中的真实流形结构,我们倡导在最小最大博弈中针对新型结构判别器对扩散生成器进行对抗训练,从而区分真实流形结构与生成的流形结构。SADM显著改进了现有的扩散Transformer(DiT),并在12个数据集的图像生成和跨域微调任务中超越了现有方法,在ImageNet上针对256x256和512x512分辨率类别条件图像生成分别建立了FID为1.58和2.11的最新最佳水平。