Leveraging multi-view diffusion models as priors for 3D optimization have alleviated the problem of 3D consistency, e.g., the Janus face problem or the content drift problem, in zero-shot text-to-3D models. However, the 3D geometric fidelity of the output remains an unresolved issue; albeit the rendered 2D views are realistic, the underlying geometry may contain errors such as unreasonable concavities. In this work, we propose CorrespondentDream, an effective method to leverage annotation-free, cross-view correspondences yielded from the diffusion U-Net to provide additional 3D prior to the NeRF optimization process. We find that these correspondences are strongly consistent with human perception, and by adopting it in our loss design, we are able to produce NeRF models with geometries that are more coherent with common sense, e.g., more smoothed object surface, yielding higher 3D fidelity. We demonstrate the efficacy of our approach through various comparative qualitative results and a solid user study.
翻译:利用多视图扩散模型作为3D优化的先验,已缓解了零样本文本到3D模型中的3D一致性问题(如Janus面问题或内容漂移问题)。然而,输出结果的3D几何保真度仍是一个未解决的问题:尽管渲染的2D视图具有真实感,但底层几何结构可能包含错误(例如不合理的凹陷)。在本工作中,我们提出CorrespondentDream,一种有效方法,利用扩散U-Net生成的无标注跨视角对应关系为NeRF优化过程提供额外3D先验。我们发现这些对应关系与人类感知高度一致,通过将其融入损失函数设计,能够生成几何结构更符合常识(例如更平滑的物体表面)的NeRF模型,从而获得更高的3D保真度。我们通过多种对比定性结果和扎实的用户研究验证了该方法的有效性。