Optical flow estimation is crucial for various applications in vision and robotics. As the difficulty of collecting ground truth optical flow in real-world scenarios, most of the existing methods of learning optical flow still adopt synthetic dataset for supervised training or utilize photometric consistency across temporally adjacent video frames to drive the unsupervised learning, where the former typically has issues of generalizability while the latter usually performs worse than the supervised ones. To tackle such challenges, we propose to leverage the geometric connection between optical flow estimation and stereo matching (based on the similarity upon finding pixel correspondences across images) to unify various real-world depth estimation datasets for generating supervised training data upon optical flow. Specifically, we turn the monocular depth datasets into stereo ones via synthesizing virtual disparity, thus leading to the flows along the horizontal direction; moreover, we introduce virtual camera motion into stereo data to produce additional flows along the vertical direction. Furthermore, we propose applying geometric augmentations on one image of an optical flow pair, encouraging the optical flow estimator to learn from more challenging cases. Lastly, as the optical flow maps under different geometric augmentations actually exhibit distinct characteristics, an auxiliary classifier which trains to identify the type of augmentation from the appearance of the flow map is utilized to further enhance the learning of the optical flow estimator. Our proposed method is general and is not tied to any particular flow estimator, where extensive experiments based on various datasets and optical flow estimation models verify its efficacy and superiority.
翻译:光流估计对于视觉和机器人领域的各类应用至关重要。由于在真实场景中收集光流真值存在困难,现有的大多数光流学习方法仍采用合成数据集进行监督训练,或利用时间相邻视频帧间的光度一致性驱动无监督学习。前者通常存在泛化性问题,而后者表现往往逊于监督方法。为应对这些挑战,本文提出利用光流估计与立体匹配(基于跨图像像素对应搜索的相似性)之间的几何关联,整合多种真实深度估计数据集,为光流生成监督训练数据。具体而言,通过合成虚拟视差将单目深度数据集转化为立体数据集,从而生成沿水平方向的光流;进一步,在立体数据中引入虚拟相机运动以产生沿垂直方向的额外光流。此外,提出对光流对中的单幅图像施加几何增广,促使光流估计器从更具挑战性的案例中学习。最后,由于不同几何增广下的光流图实际上展现出不同特征,我们利用辅助分类器(训练用于从光流图外观识别增广类型)进一步增强光流估计器的学习效果。所提方法具有通用性,不依赖于特定流估计器,基于多种数据集和光流估计模型的广泛实验验证了其有效性与优越性。