Neural Radiance Fields (NeRFs) have recently emerged as a popular option for photo-realistic object capture due to their ability to faithfully capture high-fidelity volumetric content even from handheld video input. Although much research has been devoted to efficient optimization leading to real-time training and rendering, options for interactive editing NeRFs remain limited. We present a very simple but effective neural network architecture that is fast and efficient while maintaining a low memory footprint. This architecture can be incrementally guided through user-friendly image-based edits. Our representation allows straightforward object selection via semantic feature distillation at the training stage. More importantly, we propose a local 3D-aware image context to facilitate view-consistent image editing that can then be distilled into fine-tuned NeRFs, via geometric and appearance adjustments. We evaluate our setup on a variety of examples to demonstrate appearance and geometric edits and report 10-30x speedup over concurrent work focusing on text-guided NeRF editing. Video results can be seen on our project webpage at https://proteusnerf.github.io.
翻译:神经辐射场(Neural Radiance Fields,NeRFs)近年来因其即便从手持视频输入中也能忠实捕获高保真体积内容的能力,成为照片级真实物体捕捉的热门选择。尽管已有大量研究致力于高效优化以实现实时训练和渲染,但交互式编辑NeRF的选项仍然有限。我们提出了一种非常简单但有效的神经网络架构,该架构兼具快速与高效的特点,同时保持较低的内存占用。该架构可通过用户友好的基于图像的编辑进行增量式引导。我们的表示方法通过在训练阶段进行语义特征蒸馏,实现了直观的对象选择。更重要的是,我们提出了一种局部三维感知图像上下文,以促进视角一致的图像编辑,随后通过几何与外观调整将其蒸馏至微调后的NeRF中。我们在多个示例上评估了该框架,以展示外观与几何编辑效果,并报告了相较于同期专注于文本引导NeRF编辑的工作,实现了10-30倍的速度提升。视频结果可在我们的项目页面https://proteusnerf.github.io上查看。