We propose a method for editing NeRF scenes with text-instructions. Given a NeRF of a scene and the collection of images used to reconstruct it, our method uses an image-conditioned diffusion model (InstructPix2Pix) to iteratively edit the input images while optimizing the underlying scene, resulting in an optimized 3D scene that respects the edit instruction. We demonstrate that our proposed method is able to edit large-scale, real-world scenes, and is able to accomplish more realistic, targeted edits than prior work.
翻译:我们提出一种利用文本指令编辑NeRF场景的方法。给定场景的NeRF及其重建所用的图像集合,本方法采用图像条件扩散模型(InstructPix2Pix),在优化底层场景的同时迭代编辑输入图像,最终生成符合编辑指令的优化三维场景。实验表明,该方法能编辑大规模真实场景,且相比先前工作能实现更逼真、更具针对性的编辑效果。