We propose MVDream, a multi-view diffusion model that is able to generate geometrically consistent multi-view images from a given text prompt. By leveraging image diffusion models pre-trained on large-scale web datasets and a multi-view dataset rendered from 3D assets, the resulting multi-view diffusion model can achieve both the generalizability of 2D diffusion and the consistency of 3D data. Such a model can thus be applied as a multi-view prior for 3D generation via Score Distillation Sampling, where it greatly improves the stability of existing 2D-lifting methods by solving the 3D consistency problem. Finally, we show that the multi-view diffusion model can also be fine-tuned under a few shot setting for personalized 3D generation, i.e. DreamBooth3D application, where the consistency can be maintained after learning the subject identity.
翻译:我们提出MVDream——一种能够从给定文本提示生成几何一致的多视图图像的多视图扩散模型。通过利用在大型网络数据集上预训练的2D图像扩散模型以及从三维资产渲染的多视图数据集,所得到的多视图扩散模型能够同时实现2D扩散的泛化能力与3D数据的一致性。该模型可被用作基于分数蒸馏采样的三维生成中的多视图先验,通过解决三维一致性问题极大提升现有2D提升方法的稳定性。最后,我们证明该多视图扩散模型还能在少样本场景下进行微调以实现个性化三维生成(即DreamBooth3D应用),在学习目标主体身份后仍能保持一致性。