Modern video generators produce visually compelling clips but still struggle with physical and motion consistency, limiting their use as reliable world simulators. Existing remedies often rely on external simulators, teacher models, or curated physics-focused data. We explore a complementary self-supervised direction: extracting motion cues from the unlabeled videos already used to train video diffusion models. We propose LaMo, which formulates a latent motion prior over frame-to-frame latent changes conditioned on the current latent and prompt. This prior is exposed through two lightweight readouts: a macro motion drift used during training as a Motion Drift Loss, and a learned micro motion field used during sampling as Motion Prior Guidance. Both components are plug-and-play with existing video diffusion backbones, requiring no architectural or I/O changes. On VideoPhy and VideoPhy2, LaMo improves CogVideoX backbones and outperforms recent physics-aware baselines that use external supervision. On VBench, it preserves overall generation quality while improving motion-related dimensions. These results suggest that unlabeled video contains useful motion supervision for improving physical fidelity in modern video diffusion models.
翻译:现代视频生成器能够产生视觉上引人入胜的片段,但在物理一致性和运动一致性方面仍然存在困难,限制了它们作为可靠世界模拟器的应用。现有的补救措施通常依赖外部模拟器、教师模型或精心策划的物理聚焦数据。我们探索了一种互补的自监督方向:从已经用于训练视频扩散模型的无标签视频中提取运动线索。我们提出了LaMo,它构建了一个潜运动先验,用于描述帧间潜变量的变化,该变化以当前潜变量和提示为条件。该先验通过两种轻量级读出机制实现:一种是在训练过程中用作运动漂移损失的宏运动漂移,另一种是在采样过程中用作运动先验指导的微观运动场。这两个组件均可即插即用于现有视频扩散主干网络,无需任何架构或输入输出变更。在VideoPhy和VideoPhy2数据集上,LaMo改进了CogVideoX主干网络,并优于使用外部监督的近期物理感知基线方法。在VBench上,它在保持整体生成质量的同时,改善了与运动相关的维度。这些结果表明,无标签视频中包含有用的运动监督信息,可用于提升现代视频扩散模型的物理真实性。