Benefiting from the rapid development of 2D diffusion models, 3D content creation has made significant progress recently. One promising solution involves the fine-tuning of pre-trained 2D diffusion models to harness their capacity for producing multi-view images, which are then lifted into accurate 3D models via methods like fast-NeRFs or large reconstruction models. However, as inconsistency still exists and limited generated resolution, the generation results of such methods still lack intricate textures and complex geometries. To solve this problem, we propose Magic-Boost, a multi-view conditioned diffusion model that significantly refines coarse generative results through a brief period of SDS optimization ($\sim15$min). Compared to the previous text or single image based diffusion models, Magic-Boost exhibits a robust capability to generate images with high consistency from pseudo synthesized multi-view images. It provides precise SDS guidance that well aligns with the identity of the input images, enriching the local detail in both geometry and texture of the initial generative results. Extensive experiments show Magic-Boost greatly enhances the coarse inputs and generates high-quality 3D assets with rich geometric and textural details. (Project Page: https://magic-research.github.io/magic-boost/)
翻译:受益于二维扩散模型的快速发展,三维内容创作近期取得了显著进展。一种颇具前景的方案是微调预训练的二维扩散模型,利用其生成多视角图像的能力,再通过fast-NeRF或大型重建模型等方法将其提升为精确的三维模型。然而,由于仍存在视角不一致性且生成分辨率有限,这类方法的生成结果仍缺乏精细纹理和复杂几何结构。为解决该问题,我们提出Magic-Boost——一种多视角条件扩散模型,通过短时的SDS优化(约15分钟)显著优化粗粒度的生成结果。与以往基于文本或单张图像的扩散模型相比,Magic-Boost具备从伪合成多视角图像生成高一致性图像的稳健能力。它能提供与输入图像身份高度吻合的精确SDS引导,丰富初始生成结果的几何与纹理局部细节。大量实验证明,Magic-Boost显著提升了粗粒度输入的质量,生成了具有丰富几何与纹理细节的高质量三维资产。(项目主页:https://magic-research.github.io/magic-boost/)