We present Flex4DHuman, a multi-view video diffusion model that transforms a monocular or sparse multi-view video of a dynamic subject into synchronized dense multi-view videos using only relative camera-pose conditioning. Unlike prior human-centric methods that rely on skeletons, depth maps, normals, or rendered target-view geometry, Flex4DHuman requires no explicit geometry priors and instead conditions generation through relative camera-pose positional encoding. The generated videos can be directly ingested by downstream reconstruction pipelines to create dynamic 4D Gaussian splats. Built on the Wan 2.1 1.3B text-to-video model, Flex4DHuman preserves the backbone architecture and encodes camera and view information through a five-axis positional encoding that extends spatio-temporal RoPE with view indices and continuous SE(3) relative camera geometry. A three-stage curriculum progressively trains the model for pose following, flexible reference-to-target view generation, and temporal rollout. To support temporal rollout, we train with clean historical target-view tokens. We also add multi-view captions to enable test-time text control. Combined with an off-the-shelf 4D Gaussian Splatting stage, our framework lifts monocular static-camera videos into dynamic 4D Gaussian splats. Experiments on DNA-Rendering and ActorsHQ show that Flex4DHuman surpasses prior state-of-the-art methods, while the same formulation generalizes to animal categories after mixed human-animal training. These capabilities make Flex4DHuman a practical step toward scalable 4D content creation from casual monocular videos for simulation, gaming, AR/VR, and video re-shooting.
翻译:摘要:我们提出了Flex4DHuman,一种多视角视频扩散模型,能够将动态主体的单目或稀疏多视角视频,仅通过相对相机姿态条件,转换为同步的密集多视角视频。与先前依赖骨骼、深度图、法线图或渲染目标视角几何的人体中心方法不同,Flex4DHuman无需显式几何先验,而是通过相对相机姿态的位置编码来条件化生成过程。生成的视频可直接输入下游重建流程,用于创建动态四维高斯泼溅。Flex4DHuman基于Wan 2.1 1.3B文本生成视频模型,保留了主干架构,并通过五轴位置编码对相机和视角信息进行编码,该编码将时空旋转位置编码与视角索引和连续SE(3)相对相机几何相结合。一个三阶段课程学习逐步训练模型以进行姿态跟踪、灵活的参考到目标视角生成以及时间序列展开。为支持时间序列展开,我们使用干净的历目标视角令牌进行训练。我们还添加了多视角文本描述,以实现在测试时进行文本控制。结合现成的四维高斯泼溅阶段,我们的框架将单目静态相机视频提升为动态四维高斯泼溅。在DNA-Rendering和ActorsHQ上的实验表明,Flex4DHuman超越了先前最先进的方法,而相同的公式在混合人-动物训练后,可推广到动物类别。这些能力使Flex4DHuman成为从随意拍摄的单目视频实现可扩展四维内容生成的实际步骤,适用于仿真、游戏、增强现实/虚拟现实以及视频重拍。