Numerous diffusion models have recently been applied to image synthesis and editing. However, editing 3D scenes is still in its early stages. It poses various challenges, such as the requirement to design specific methods for different editing types, retraining new models for various 3D scenes, and the absence of convenient human interaction during editing. To tackle these issues, we introduce a text-driven editing method, termed DN2N, which allows for the direct acquisition of a NeRF model with universal editing capabilities, eliminating the requirement for retraining. Our method employs off-the-shelf text-based editing models of 2D images to modify the 3D scene images, followed by a filtering process to discard poorly edited images that disrupt 3D consistency. We then consider the remaining inconsistency as a problem of removing noise perturbation, which can be solved by generating training data with similar perturbation characteristics for training. We further propose cross-view regularization terms to help the generalized NeRF model mitigate these perturbations. Our text-driven method allows users to edit a 3D scene with their desired description, which is more friendly, intuitive, and practical than prior works. Empirical results show that our method achieves multiple editing types, including but not limited to appearance editing, weather transition, material changing, and style transfer. Most importantly, our method generalizes well with editing abilities shared among a set of model parameters without requiring a customized editing model for some specific scenes, thus inferring novel views with editing effects directly from user input. The project website is available at https://sk-fun.fun/DN2N
翻译:近年来,众多扩散模型已被应用于图像合成与编辑领域。然而,3D场景的编辑仍处于早期阶段,面临诸多挑战:需针对不同编辑类型设计特定方法、为各类3D场景重新训练新模型、以及编辑过程中缺乏便捷的人机交互。为解决这些问题,我们提出一种名为DN2N的文本驱动编辑方法,该方法可直接获取具备通用编辑能力的NeRF模型,无需重新训练。我们的方法利用现成的基于文本的2D图像编辑模型修改3D场景图像,随后通过过滤流程剔除破坏3D一致性的低质量编辑图像。我们将剩余的不一致性视为去噪扰动问题,通过生成具有类似扰动特征的训练数据进行解决。进一步地,我们提出跨视角正则化项,帮助泛化NeRF模型缓解这些扰动。与先前工作相比,本方法允许用户通过自然语言描述编辑3D场景,更加友好、直观且实用。实验结果表明,本方法支持多种编辑类型,包括但不限于外观编辑、天气转换、材质更换与风格迁移。最重要的是,本方法通过一组共享编辑能力的模型参数实现良好泛化,无需为特定场景定制编辑模型,可直接根据用户输入生成带有编辑效果的新视角图像。项目网站见https://sk-fun.fun/DN2N