Text-to-3D generation from a single-view image is a popular but challenging task in 3D vision. Although numerous methods have been proposed, existing works still suffer from the inconsistency issues, including 1) semantic inconsistency, 2) geometric inconsistency, and 3) saturation inconsistency, resulting in distorted, overfitted, and over-saturated generations. In light of the above issues, we present Consist3D, a three-stage framework Chasing for semantic-, geometric-, and saturation-Consistent Text-to-3D generation from a single image, in which the first two stages aim to learn parameterized consistency tokens, and the last stage is for optimization. Specifically, the semantic encoding stage learns a token independent of views and estimations, promoting semantic consistency and robustness. Meanwhile, the geometric encoding stage learns another token with comprehensive geometry and reconstruction constraints under novel-view estimations, reducing overfitting and encouraging geometric consistency. Finally, the optimization stage benefits from the semantic and geometric tokens, allowing a low classifier-free guidance scale and therefore preventing oversaturation. Experimental results demonstrate that Consist3D produces more consistent, faithful, and photo-realistic 3D assets compared to previous state-of-the-art methods. Furthermore, Consist3D also allows background and object editing through text prompts.
翻译:从单视角图像生成文本引导的3D内容,是三维视觉领域一项热门但颇具挑战的任务。尽管已有大量方法被提出,现有工作仍面临不一致性问题,包括:1)语义不一致性、2)几何不一致性、3)饱和度不一致性,导致生成结果出现扭曲、过拟合和过饱和现象。针对上述问题,我们提出Consist3D——一个三阶段框架,旨在从单张图像实现语义、几何与饱和度一致的文本到3D生成:前两阶段分别学习参数化的一致性表征令牌,第三阶段进行优化。具体而言,语义编码阶段学习一个与视角和估计无关的令牌,促进语义一致性与鲁棒性;几何编码阶段在新视角估计下,通过全面的几何与重建约束学习另一个令牌,减少过拟合并增强几何一致性;优化阶段得益于语义与几何令牌,可降低无分类器引导尺度,从而避免过饱和。实验结果表明,与现有最先进方法相比,Consist3D能生成更一致、保真且逼真的3D资产。此外,Consist3D还支持通过文本提示对背景和物体进行编辑。