Recent research endeavors have shown that combining neural radiance fields (NeRFs) with pre-trained diffusion models holds great potential for text-to-3D generation.However, a hurdle is that they often encounter guidance collapse when rendering complex scenes from multi-object texts. Because the text-to-image diffusion models are inherently unconstrained, making them less competent to accurately associate object semantics with specific 3D structures. To address this issue, we propose a novel framework, dubbed CompoNeRF, that explicitly incorporates an editable 3D scene layout to provide effective guidance at the single object (i.e., local) and whole scene (i.e., global) levels. Firstly, we interpret the multi-object text as an editable 3D scene layout containing multiple local NeRFs associated with the object-specific 3D box coordinates and text prompt, which can be easily collected from users. Then, we introduce a global MLP to calibrate the compositional latent features from local NeRFs, which surprisingly improves the view consistency across different local NeRFs. Lastly, we apply the text guidance on global and local levels through their corresponding views to avoid guidance ambiguity. This way, our CompoNeRF allows for flexible scene editing and re-composition of trained local NeRFs into a new scene by manipulating the 3D layout or text prompt. Leveraging the open-source Stable Diffusion model, our CompoNeRF can generate faithful and editable text-to-3D results while opening a potential direction for text-guided multi-object composition via the editable 3D scene layout.


翻译:近期研究尝试将神经辐射场(NeRFs)与预训练的扩散模型相结合,在文本到三维生成领域展现出巨大潜力。然而,一个关键障碍在于:当从多对象文本渲染复杂场景时,此类方法常遭遇引导坍缩问题。这是因为文本到图像的扩散模型本质不受约束,难以准确将对象语义与特定三维结构相关联。为解决该问题,我们提出名为CompoNeRF的新型框架,该框架显式融合可编辑三维场景布局,在单对象(即局部)和全场景(即全局)层面提供有效引导。首先,我们将多对象文本解析为包含多个局部NeRF的可编辑三维场景布局,每个局部NeRF关联对象特定的三维包围盒坐标与文本提示(这些信息可由用户便捷提供)。随后,引入全局MLP对源自各局部NeRF的组合隐特征进行校准,这一设计显著提升了不同局部NeRF间的视角一致性。最终,我们通过全局与局部视角分别施加文本引导,以避免引导歧义。通过上述设计,CompoNeRF可支持灵活的场景编辑——通过操纵三维布局或文本提示,将已训练的各局部NeRF重组至新场景。基于开源Stable Diffusion模型,CompoNeRF不仅能生成忠实且可编辑的文本到三维结果,更开辟了通过可编辑三维场景布局实现文本引导多对象组合的潜在研究方向。

0
下载
关闭预览

相关内容

【CVPR2023】NS3D:3D对象和关系的神经符号Grounding
专知会员服务
23+阅读 · 2023年3月26日
【AAAI2023】用于复杂场景图像合成的特征金字塔扩散模型
专知会员服务
11+阅读 · 2021年8月11日
专知会员服务
74+阅读 · 2021年5月28日
EMNLP 2022 | 统一指代性表达的生成和理解
PaperWeekly
1+阅读 · 2022年11月8日
ECCV 2022 | 底层视觉新任务:Blind Image Decomposition
【泡泡一分钟】DS-SLAM: 动态环境下的语义视觉SLAM
泡泡机器人SLAM
23+阅读 · 2019年1月18日
MoCoGAN 分解运动和内容的视频生成
CreateAMind
18+阅读 · 2017年10月21日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2011年12月31日
国家自然科学基金
0+阅读 · 2009年12月31日
国家自然科学基金
4+阅读 · 2009年12月31日
Arxiv
0+阅读 · 2023年5月14日
VIP会员
最新内容
深入Project Maven:为何人工智能在战场上依然失灵
锻造未来士兵:外骨骼、基因工程与赛博格
专知会员服务
6+阅读 · 7月19日
《无人机蜂群通信技术研究》50页
专知会员服务
7+阅读 · 7月19日
战力倍增器:自主武器系统与乌克兰及加沙冲突
人工智能赋能战场情报:提速决策进程
专知会员服务
5+阅读 · 7月17日
《拥抱新兴技术:面向未来军官的教育革新》
专知会员服务
8+阅读 · 7月17日
相关VIP内容
【CVPR2023】NS3D:3D对象和关系的神经符号Grounding
专知会员服务
23+阅读 · 2023年3月26日
【AAAI2023】用于复杂场景图像合成的特征金字塔扩散模型
专知会员服务
11+阅读 · 2021年8月11日
专知会员服务
74+阅读 · 2021年5月28日
相关基金
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2011年12月31日
国家自然科学基金
0+阅读 · 2009年12月31日
国家自然科学基金
4+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员