Reconstructing dynamic 3D scenes from 2D images and generating diverse views over time is challenging due to scene complexity and temporal dynamics. Despite advancements in neural implicit models, limitations persist: (i) Inadequate Scene Structure: Existing methods struggle to reveal the spatial and temporal structure of dynamic scenes from directly learning the complex 6D plenoptic function. (ii) Scaling Deformation Modeling: Explicitly modeling scene element deformation becomes impractical for complex dynamics. To address these issues, we consider the spacetime as an entirety and propose to approximate the underlying spatio-temporal 4D volume of a dynamic scene by optimizing a collection of 4D primitives, with explicit geometry and appearance modeling. Learning to optimize the 4D primitives enables us to synthesize novel views at any desired time with our tailored rendering routine. Our model is conceptually simple, consisting of a 4D Gaussian parameterized by anisotropic ellipses that can rotate arbitrarily in space and time, as well as view-dependent and time-evolved appearance represented by the coefficient of 4D spherindrical harmonics. This approach offers simplicity, flexibility for variable-length video and end-to-end training, and efficient real-time rendering, making it suitable for capturing complex dynamic scene motions. Experiments across various benchmarks, including monocular and multi-view scenarios, demonstrate our 4DGS model's superior visual quality and efficiency.
翻译:从二维图像重建动态三维场景并随时间生成多样视角是一项具有挑战性的任务,原因在于场景的复杂性和时间动态性。尽管神经隐式模型取得了进展,但仍存在局限性:(i)场景结构不充分:现有方法难以从直接学习复杂的六维全光函数中揭示动态场景的时空结构。(ii)形变建模的扩展性:显式建模场景元素形变对于复杂动态而言变得不切实际。为解决这些问题,我们将时空视为一个整体,提出通过优化一组具有显式几何与外观建模的四维基元来近似动态场景的底层时空四维体。通过学习优化这些四维基元,我们能够利用定制的渲染流程在任意所需时间合成新视角。我们的模型在概念上简洁,由四维高斯函数参数化,该函数基于可在空间和时间上任意旋转的各向异性椭圆,以及由四维球柱谐波系数表示的视角相关且随时间演进的外观。该方法具有简洁性、对可变长度视频的灵活性和端到端训练能力,以及高效的实时渲染,使其适用于捕捉复杂动态场景运动。在包括单目和多视角场景在内的多个基准测试上的实验,证明了我们4DGS模型卓越的视觉质量和效率。