Dynamic scene reconstruction and novel view synthesis are fundamental to next-generation visual intelligence applications such as virtual reality, robotics, and digital twins. However, high-fidelity reconstruction of complex, time-varying scenes from arbitrary viewpoints remains a significant challenge. Existing dynamic 3DGS methods suffer from computational inefficiency, since they model all Gaussians as dynamic components. While recent decomposition-based approaches address this issue, they still struggle with degraded reconstruction quality and prolonged training time. To mitigate these limitations, we propose a novel dynamic reconstruction framework built upon an efficient static-dynamic decomposition strategy using a Feed-Forward Gaussian Splatting encoder and an optical flow model. By eliminating redundant computations on static regions, our method achieves state-of-the-art performance, outperforming existing baselines across rendering quality, training and rendering speed, and storage efficiency. Notably, on the Neural 3D dataset, our framework requires only 10 minutes for training and achieves a rendering speed of over 700 FPS on a single NVIDIA RTX 5090 GPU at resolution of 1352x1014. Furthermore, our decomposition strategy eliminates the need for COLMAP preprocessing and enables deterministic initialization, thereby enhancing both efficiency and reproducibility.
翻译:动态场景重建与新视角合成是实现虚拟现实、机器人技术和数字孪生等下一代视觉智能应用的基础。然而,从任意视角对复杂时变场景进行高保真重建仍是一项重大挑战。现有动态三维高斯泼溅方法将全部高斯组件均建模为动态分量,导致计算效率低下。尽管近期基于分解的策略试图解决此问题,但其重建质量下降与训练耗时延长的问题依然存在。为缓解上述局限,我们提出一种新型动态重建框架,该框架基于高效的静态-动态分解策略,集成前馈式高斯泼溅编码器与光流模型。通过消除静态区域的冗余计算,本方法在渲染质量、训练与渲染速度、存储效率等指标上全面超越现有基线,达到最先进水平。值得注意的是,在Neural 3D数据集上,本框架仅需10分钟训练,即可在单块NVIDIA RTX 5090 GPU上以1352×1014分辨率实现超过700 FPS的渲染速度。此外,该分解策略消除了对COLMAP预处理的依赖,并支持确定性初始化,从而同时提升效率与可复现性。