Pairwise matching cost aggregation is a crucial step for modern learning-based Multi-view Stereo (MVS). Prior works adopt an early aggregation scheme, which adds up pairwise costs into an intermediate cost. However, we analyze that this process can degrade informative pairwise matchings, thereby blocking the depth network from fully utilizing the original geometric matching cues. To address this challenge, we present a late aggregation approach that allows for aggregating pairwise costs throughout the network feed-forward process, achieving accurate estimations with only minor changes of the plain CasMVSNet. Instead of building an intermediate cost by weighted sum, late aggregation preserves all pairwise costs along a distinct view channel. This enables the succeeding depth network to fully utilize the crucial geometric cues without loss of cost fidelity. Grounded in the new aggregation scheme, we propose further techniques addressing view order dependence inside the preserved cost, handling flexible testing views, and improving the depth filtering process. Despite its technical simplicity, our method improves significantly upon the baseline cascade-based approach, achieving comparable results with state-of-the-art methods with favorable computation overhead.
翻译:成对匹配代价聚合是现代基于学习的多视图立体(MVS)的关键步骤。现有方法采用早期聚合策略,将成对代价累加为中间代价值。然而,我们分析发现该过程会削弱信息丰富的成对匹配关系,从而阻碍深度网络充分利用原始的几何匹配线索。为应对这一挑战,我们提出一种晚期聚合方法,允许在网络前馈过程中聚合成对代价,仅需对普通CasMVSNet进行微小改动即可实现精确估计。晚期聚合不通过加权求和构建中间代价,而是沿独立视角通道保留所有成对代价,使后续深度网络能在不损失代价保真度的情况下充分利用关键几何线索。基于这一新型聚合方案,我们进一步提出解决保留代价中的视角顺序依赖性、处理灵活测试视角以及改进深度滤波过程的技术。尽管该方法技术简洁,但在级联基线方法基础上实现了显著提升,以可控的计算开销达到了与最先进方法相当的性能。