Diffusion models currently achieve state-of-the-art performance for both conditional and unconditional image generation. However, so far, image diffusion models do not support tasks required for 3D understanding, such as view-consistent 3D generation or single-view object reconstruction. In this paper, we present RenderDiffusion, the first diffusion model for 3D generation and inference, trained using only monocular 2D supervision. Central to our method is a novel image denoising architecture that generates and renders an intermediate three-dimensional representation of a scene in each denoising step. This enforces a strong inductive structure within the diffusion process, providing a 3D consistent representation while only requiring 2D supervision. The resulting 3D representation can be rendered from any view. We evaluate RenderDiffusion on FFHQ, AFHQ, ShapeNet and CLEVR datasets, showing competitive performance for generation of 3D scenes and inference of 3D scenes from 2D images. Additionally, our diffusion-based approach allows us to use 2D inpainting to edit 3D scenes.
翻译:扩散模型当前在有条件与无条件图像生成任务中均达到最优性能。然而,现有图像扩散模型尚不支持三维理解所需的任务,例如视角一致性三维生成或单视图物体重建。本文提出RenderDiffusion——首个仅需单目二维监督即可实现三维生成与推理的扩散模型。该方法的核心是一种新型图像去噪架构,该架构在每个去噪步骤中生成并渲染场景的中间三维表示。这为扩散过程施加了强归纳结构,在仅需二维监督的同时提供三维一致表示。生成的三维表示可从任意视角进行渲染。我们在FFHQ、AFHQ、ShapeNet及CLEVR数据集上评估RenderDiffusion,结果表明其在三维场景生成与二维图像三维场景推理方面具有竞争力。此外,基于扩散的方法使我们能够利用二维图像修复技术编辑三维场景。