Diffusion models have recently become the de-facto approach for generative modeling in the 2D domain. However, extending diffusion models to 3D is challenging due to the difficulties in acquiring 3D ground truth data for training. On the other hand, 3D GANs that integrate implicit 3D representations into GANs have shown remarkable 3D-aware generation when trained only on single-view image datasets. However, 3D GANs do not provide straightforward ways to precisely control image synthesis. To address these challenges, We present Control3Diff, a 3D diffusion model that combines the strengths of diffusion models and 3D GANs for versatile, controllable 3D-aware image synthesis for single-view datasets. Control3Diff explicitly models the underlying latent distribution (optionally conditioned on external inputs), thus enabling direct control during the diffusion process. Moreover, our approach is general and applicable to any type of controlling input, allowing us to train it with the same diffusion objective without any auxiliary supervision. We validate the efficacy of Control3Diff on standard image generation benchmarks, including FFHQ, AFHQ, and ShapeNet, using various conditioning inputs such as images, sketches, and text prompts. Please see the project website (\url{https://jiataogu.me/control3diff}) for video comparisons.
翻译:扩散模型近来已成为2D领域生成建模的事实标准。然而,将扩散模型扩展到3D面临挑战,主要源于获取3D真实训练数据的困难。另一方面,将隐式3D表征融入生成对抗网络(GAN)的3D GAN,在仅依靠单视图图像数据集训练时已展现出显著的3D感知生成能力。但3D GAN无法提供直接途径来精确控制图像合成。为应对这些挑战,我们提出了Control3Diff——一种融合扩散模型与3D GAN优势的3D扩散模型,能够针对单视图数据集实现灵活可控的3D感知图像合成。Control3Diff显式建模底层潜在分布(可选条件约束于外部输入),从而在扩散过程中实现直接控制。此外,我们的方法具有通用性,适用于任意类型的控制输入,可仅使用相同的扩散目标进行训练而无需辅助监督。我们在FFHQ、AFHQ和ShapeNet等标准图像生成基准上,利用图像、草图及文本提示等多种条件输入验证了Control3Diff的有效性。视频对比请参见项目网站(\url{https://jiataogu.me/control3diff})。