Reasoning the 3D structure of a non-rigid dynamic scene from a single moving camera is an under-constrained problem. Inspired by the remarkable progress of neural radiance fields (NeRFs) in photo-realistic novel view synthesis of static scenes, extensions have been proposed for dynamic settings. These methods heavily rely on neural priors in order to regularize the problem. In this work, we take a step back and reinvestigate how current implementations may entail deleterious effects, including limited expressiveness, entanglement of light and density fields, and sub-optimal motion localization. As a remedy, we advocate for a bridge between classic non-rigid-structure-from-motion (\nrsfm) and NeRF, enabling the well-studied priors of the former to constrain the latter. To this end, we propose a framework that factorizes time and space by formulating a scene as a composition of bandlimited, high-dimensional signals. We demonstrate compelling results across complex dynamic scenes that involve changes in lighting, texture and long-range dynamics.
翻译:从单个移动相机推理非刚性动态场景的三维结构是一个欠约束问题。受神经辐射场(NeRF)在静态场景逼真新视角合成中取得显著进展的启发,研究者提出了针对动态场景的扩展方法。这些方法严重依赖神经先验来正则化问题。本研究重新审视当前实现可能带来的弊端,包括表达力受限、光场与密度场纠缠以及运动定位欠优。作为解决方案,我们主张在经典的非刚性运动恢复结构(NRSfM)与NeRF之间建立桥梁,利用前者经过充分研究的先验约束后者。为此,我们提出一个框架,通过将场景表示为带限高维信号的组合来实现时间与空间分解。我们在涉及光照变化、纹理变化及长程动态的复杂动态场景上展示了令人信服的结果。