Traditional multi-view stereo (MVS) methods rely heavily on photometric and geometric consistency constraints, but newer machine learning-based MVS methods check geometric consistency across multiple source views only as a post-processing step. In this paper, we present a novel approach that explicitly encourages geometric consistency of reference view depth maps across multiple source views at different scales during learning (see Fig. 1). We find that adding this geometric consistency loss significantly accelerates learning by explicitly penalizing geometrically inconsistent pixels, reducing the training iteration requirements to nearly half that of other MVS methods. Our extensive experiments show that our approach achieves a new state-of-the-art on the DTU and BlendedMVS datasets, and competitive results on the Tanks and Temples benchmark. To the best of our knowledge, GC-MVSNet is the first attempt to enforce multi-view, multi-scale geometric consistency during learning.
翻译:传统多视图立体(MVS)方法严重依赖光度一致性和几何一致性约束,但基于机器学习的现代MVS方法仅在后期处理步骤中检查多源视图的几何一致性。本文提出一种新颖方法,在学习过程中显式地鼓励参考视图深度图在不同尺度上跨多源视图保持几何一致性(见图1)。我们发现,通过显式惩罚几何不一致的像素,这种几何一致性损失显著加速了学习进程,使训练迭代次数减少至其他MVS方法的一半左右。大量实验表明,我们的方法在DTU和BlendedMVS数据集上达到了新的最优水平,在Tanks and Temples基准上获得了具有竞争力的结果。据我们所知,GC-MVSNet是首个在学习过程中强制实现多视角、多尺度几何一致性的尝试。