We present a new test-time optimization method for estimating dense and long-range motion from a video sequence. Prior optical flow or particle video tracking algorithms typically operate within limited temporal windows, struggling to track through occlusions and maintain global consistency of estimated motion trajectories. We propose a complete and globally consistent motion representation, dubbed OmniMotion, that allows for accurate, full-length motion estimation of every pixel in a video. OmniMotion represents a video using a quasi-3D canonical volume and performs pixel-wise tracking via bijections between local and canonical space. This representation allows us to ensure global consistency, track through occlusions, and model any combination of camera and object motion. Extensive evaluations on the TAP-Vid benchmark and real-world footage show that our approach outperforms prior state-of-the-art methods by a large margin both quantitatively and qualitatively. See our project page for more results: http://omnimotion.github.io/
翻译:我们提出一种新的测试时优化方法,用于从视频序列中估计密集且长程的运动。先前的光流或粒子视频追踪算法通常在有限的时间窗口内运行,难以追踪遮挡区域并维持估计运动轨迹的全局一致性。我们提出一种完整且全局一致的运动表示,称为OmniMotion,能够对视频中每个像素进行精确的全长运动估计。OmniMotion利用准三维的规范体积表示视频,并通过局部空间与规范空间之间的双射实现逐像素追踪。这种表示使我们能够确保全局一致性、追踪遮挡区域,并建模摄像机与物体的任意组合运动。在TAP-Vid基准测试及真实世界素材上的广泛评估表明,我们的方法在定量和定性上均大幅超越先前的最优方法。更多结果请见项目页面:http://omnimotion.github.io/