We present an approach to modeling an image-space prior on scene motion. Our prior is learned from a collection of motion trajectories extracted from real video sequences depicting natural, oscillatory dynamics such as trees, flowers, candles, and clothes swaying in the wind. We model this dense, long-term motion prior in the Fourier domain:given a single image, our trained model uses a frequency-coordinated diffusion sampling process to predict a spectral volume, which can be converted into a motion texture that spans an entire video. Along with an image-based rendering module, these trajectories can be used for a number of downstream applications, such as turning still images into seamlessly looping videos, or allowing users to realistically interact with objects in real pictures by interpreting the spectral volumes as image-space modal bases, which approximate object dynamics.
翻译:我们提出一种对场景运动的图像空间先验进行建模的方法。该先验从真实视频序列中提取的运动轨迹集合中学习,这些视频展现自然的振荡动态,如树木、花朵、蜡烛和衣物在风中摇曳。我们在傅里叶域中建模这种密集的长期运动先验:给定单张图像,我们的训练模型采用频率协调扩散采样过程来预测一个频谱体积,可将其转换为覆盖整个视频的运动纹理。结合基于图像的渲染模块,这些轨迹可用于多种下游应用,例如将静态图像转化为无缝循环视频,或通过将频谱体积解释为近似物体动态的图像空间模态基,使用户能够与真实图片中的物体进行逼真交互。