Dynamic scene reconstruction from monocular video remains a fundamental challenge in computer vision. Existing feed-forward methods predict 3D Gaussians pixel-wise for each frame, suffering from duplicated Gaussians and view-dependent biases that hinder effective learning of scene motion. We present C4G, a feed-forward 4D reconstruction framework built upon a compact set of timestamp-conditioned learnable Gaussian query tokens. Each token aggregates corresponding features across the full temporal context and decodes a 3D Gaussian whose position is modulated by the target timestamp, enabling globally coherent motion modeling without per-scene optimization. To capture fine-grained details, we further introduce a video diffusion model-based rendering enhancement module. Since our framework effectively aggregates features into Gaussians, we extend this capability to feature lifting, producing a 4D feature field that supports point tracking and dynamic scene understanding. C4G achieves strong novel-view synthesis performance using significantly fewer Gaussians and without requiring camera poses, while exhibiting stronger motion modeling and robustness to large temporal gaps.
翻译:从单目视频进行动态场景重建始终是计算机视觉领域的核心挑战。现有前馈方法逐帧逐像素预测3D高斯体,存在高斯体重复和视角依赖偏差问题,阻碍了场景运动的有效学习。我们提出C4G——一种基于紧凑型时间戳条件可学习高斯查询令牌的前馈式4D重建框架。每个令牌聚合完整时间上下文中的对应特征,并解码出一个位置受目标时间戳调制的3D高斯体,从而无需逐场景优化即可实现全局一致性运动建模。为捕捉细节,我们进一步引入基于视频扩散模型的渲染增强模块。由于本框架有效将特征聚合至高斯体,我们将该能力拓展至特征提升,构建支持点跟踪与动态场景理解的4D特征场。C4G在显著减少高斯体数量且无需相机姿态的条件下,实现了优异的新视角合成性能,并展现出更强的运动建模能力与对大时间间隔的鲁棒性。