Recent research endeavors have shown that combining neural radiance fields (NeRFs) with pre-trained diffusion models holds great potential for text-to-3D generation.However, a hurdle is that they often encounter guidance collapse when rendering complex scenes from multi-object texts. Because the text-to-image diffusion models are inherently unconstrained, making them less competent to accurately associate object semantics with specific 3D structures. To address this issue, we propose a novel framework, dubbed CompoNeRF, that explicitly incorporates an editable 3D scene layout to provide effective guidance at the single object (i.e., local) and whole scene (i.e., global) levels. Firstly, we interpret the multi-object text as an editable 3D scene layout containing multiple local NeRFs associated with the object-specific 3D box coordinates and text prompt, which can be easily collected from users. Then, we introduce a global MLP to calibrate the compositional latent features from local NeRFs, which surprisingly improves the view consistency across different local NeRFs. Lastly, we apply the text guidance on global and local levels through their corresponding views to avoid guidance ambiguity. This way, our CompoNeRF allows for flexible scene editing and re-composition of trained local NeRFs into a new scene by manipulating the 3D layout or text prompt. Leveraging the open-source Stable Diffusion model, our CompoNeRF can generate faithful and editable text-to-3D results while opening a potential direction for text-guided multi-object composition via the editable 3D scene layout.
翻译:近期研究尝试将神经辐射场(NeRFs)与预训练的扩散模型相结合,在文本到三维生成领域展现出巨大潜力。然而,一个关键障碍在于:当从多对象文本渲染复杂场景时,此类方法常遭遇引导坍缩问题。这是因为文本到图像的扩散模型本质不受约束,难以准确将对象语义与特定三维结构相关联。为解决该问题,我们提出名为CompoNeRF的新型框架,该框架显式融合可编辑三维场景布局,在单对象(即局部)和全场景(即全局)层面提供有效引导。首先,我们将多对象文本解析为包含多个局部NeRF的可编辑三维场景布局,每个局部NeRF关联对象特定的三维包围盒坐标与文本提示(这些信息可由用户便捷提供)。随后,引入全局MLP对源自各局部NeRF的组合隐特征进行校准,这一设计显著提升了不同局部NeRF间的视角一致性。最终,我们通过全局与局部视角分别施加文本引导,以避免引导歧义。通过上述设计,CompoNeRF可支持灵活的场景编辑——通过操纵三维布局或文本提示,将已训练的各局部NeRF重组至新场景。基于开源Stable Diffusion模型,CompoNeRF不仅能生成忠实且可编辑的文本到三维结果,更开辟了通过可编辑三维场景布局实现文本引导多对象组合的潜在研究方向。