For visual estimation of optical flow, a crucial function for many vision tasks, unsupervised learning, using the supervision of view synthesis has emerged as a promising alternative to supervised methods, since ground-truth flow is not readily available in many cases. However, unsupervised learning is likely to be unstable when pixel tracking is lost due to occlusion and motion blur, or the pixel matching is impaired due to variation in image content and spatial structure over time. In natural environments, dynamic occlusion or object variation is a relatively slow temporal process spanning several frames. We, therefore, explore the optical flow estimation from multiple-frame sequences of dynamic scenes, whereas most of the existing unsupervised approaches are based on temporal static models. We handle the unsupervised optical flow estimation with a temporal dynamic model by introducing a spatial-temporal dual recurrent block based on the predictive coding structure, which feeds the previous high-level motion prior to the current optical flow estimator. Assuming temporal smoothness of optical flow, we use motion priors of the adjacent frames to provide more reliable supervision of the occluded regions. To grasp the essence of challenging scenes, we simulate various scenarios across long sequences, including dynamic occlusion, content variation, and spatial variation, and adopt self-supervised distillation to make the model understand the object's motion patterns in a prolonged dynamic environment. Experiments on KITTI 2012, KITTI 2015, Sintel Clean, and Sintel Final datasets demonstrate the effectiveness of our methods on unsupervised optical flow estimation. The proposal achieves state-of-the-art performance with advantages in memory overhead.
翻译:针对视觉光流估计这一众多视觉任务的关键功能,利用视图合成监督的无监督学习已成为有监督方法的有前途替代方案,因为许多情况下无法直接获取光流真值。然而,当像素跟踪因遮挡和运动模糊而丢失,或像素匹配因图像内容与空间结构随时间变化而受损时,无监督学习可能不稳定。在自然环境中,动态遮挡或物体变化是一个跨越数帧的相对缓慢的时间过程。因此,我们探索从动态场景的多帧序列中估计光流,而现有大多数无监督方法基于时间静态模型。我们通过引入基于预测编码结构的时空双递归模块,将先前的高层运动先验馈入当前光流估计器,从而利用时间动态模型处理无监督光流估计。基于光流的时间平滑性假设,我们使用相邻帧的运动先验为遮挡区域提供更可靠的监督。为把握挑战性场景的本质,我们模拟了长序列中的多种场景,包括动态遮挡、内容变化和空间变化,并采用自监督蒸馏使模型理解长时间动态环境中物体的运动模式。在KITTI 2012、KITTI 2015、Sintel Clean和Sintel Final数据集上的实验证明了我们方法在无监督光流估计上的有效性。该方案以内存开销优势实现了最先进的性能。