We introduce a novel framework for training deep stereo networks effortlessly and without any ground-truth. By leveraging state-of-the-art neural rendering solutions, we generate stereo training data from image sequences collected with a single handheld camera. On top of them, a NeRF-supervised training procedure is carried out, from which we exploit rendered stereo triplets to compensate for occlusions and depth maps as proxy labels. This results in stereo networks capable of predicting sharp and detailed disparity maps. Experimental results show that models trained under this regime yield a 30-40% improvement over existing self-supervised methods on the challenging Middlebury dataset, filling the gap to supervised models and, most times, outperforming them at zero-shot generalization.
翻译:我们提出了一种新颖的框架,无需任何真实标注数据即可轻松训练深度立体匹配网络。通过利用最先进的神经渲染解决方案,我们从单手持相机采集的图像序列中生成立体训练数据。在此基础上,执行一种NeRF监督的训练流程,利用渲染得到的立体三元组来补偿遮挡,并将深度图作为代理标签。这使得立体匹配网络能够预测清晰且细节丰富的视差图。实验结果表明,在具有挑战性的Middlebury数据集上,基于该训练策略的模型相较于现有自监督方法实现了30-40%的性能提升,缩小了与监督模型的差距,并且在零样本泛化场景中多数情况下超越了监督模型。