While current multi-frame restoration methods combine information from multiple input images using 2D alignment techniques, recent advances in novel view synthesis are paving the way for a new paradigm relying on volumetric scene representations. In this work, we introduce the first 3D-based multi-frame denoising method that significantly outperforms its 2D-based counterparts with lower computational requirements. Our method extends the multiplane image (MPI) framework for novel view synthesis by introducing a learnable encoder-renderer pair manipulating multiplane representations in feature space. The encoder fuses information across views and operates in a depth-wise manner while the renderer fuses information across depths and operates in a view-wise manner. The two modules are trained end-to-end and learn to separate depths in an unsupervised way, giving rise to Multiplane Feature (MPF) representations. Experiments on the Spaces and Real Forward-Facing datasets as well as on raw burst data validate our approach for view synthesis, multi-frame denoising, and view synthesis under noisy conditions.
翻译:尽管当前的多帧恢复方法利用2D对齐技术结合多个输入图像的信息,但近期视图合成领域的进展正推动一种依赖体积场景表示的新范式。本文首次提出基于3D的多帧去噪方法,其性能显著超越基于2D的同类方法,且计算需求更低。我们的方法扩展了用于新型视图合成的多平面图像(MPI)框架,通过引入可学习的编码器-渲染器对,在特征空间中操控多平面表示。编码器以深度方向融合跨视图信息,而渲染器以视图方向融合跨深度信息。这两个模块以端到端方式训练,并通过无监督方式学习深度分离,从而形成多平面特征(MPF)表示。在Spaces和Real Forward-Facing数据集以及原始突发数据上的实验验证了该方法在视图合成、多帧去噪及噪声条件下视图合成中的有效性。