We explore a new class of diffusion models based on the transformer architecture. We train latent diffusion models of images, replacing the commonly-used U-Net backbone with a transformer that operates on latent patches. We analyze the scalability of our Diffusion Transformers (DiTs) through the lens of forward pass complexity as measured by Gflops. We find that DiTs with higher Gflops -- through increased transformer depth/width or increased number of input tokens -- consistently have lower FID. In addition to possessing good scalability properties, our largest DiT-XL/2 models outperform all prior diffusion models on the class-conditional ImageNet 512x512 and 256x256 benchmarks, achieving a state-of-the-art FID of 2.27 on the latter.
翻译:本文探索了一类基于Transformer架构的新型扩散模型。我们训练图像的潜在扩散模型,用操作潜在图像块的Transformer取代常用的U-Net骨干网络。我们通过前向传播复杂度(以Gflops衡量)的角度分析扩散Transformer(DiTs)的可扩展性。研究发现,具有更高Gflops的DiT(通过增加Transformer深度/宽度或增加输入令牌数量)始终能获得更低的FID。除良好的可扩展性外,我们最大的DiT-XL/2模型在类别条件ImageNet 512x512和256x256基准测试中均超越所有先前扩散模型,在256x256数据上取得了2.27的最优FID。