Multi-frame depth estimation generally achieves high accuracy relying on the multi-view geometric consistency. When applied in dynamic scenes, e.g., autonomous driving, this consistency is usually violated in the dynamic areas, leading to corrupted estimations. Many multi-frame methods handle dynamic areas by identifying them with explicit masks and compensating the multi-view cues with monocular cues represented as local monocular depth or features. The improvements are limited due to the uncontrolled quality of the masks and the underutilized benefits of the fusion of the two types of cues. In this paper, we propose a novel method to learn to fuse the multi-view and monocular cues encoded as volumes without needing the heuristically crafted masks. As unveiled in our analyses, the multi-view cues capture more accurate geometric information in static areas, and the monocular cues capture more useful contexts in dynamic areas. To let the geometric perception learned from multi-view cues in static areas propagate to the monocular representation in dynamic areas and let monocular cues enhance the representation of multi-view cost volume, we propose a cross-cue fusion (CCF) module, which includes the cross-cue attention (CCA) to encode the spatially non-local relative intra-relations from each source to enhance the representation of the other. Experiments on real-world datasets prove the significant effectiveness and generalization ability of the proposed method.
翻译:多帧深度估计通常依赖于多视角几何一致性来获得高精度。当应用于动态场景(如自动驾驶)时,这种一致性在动态区域往往被破坏,导致估计结果出现扭曲。许多多帧方法通过显式掩码识别动态区域,并用局部单目深度或特征等单目线索补偿多视角线索,但由于掩码质量不可控且未能充分利用两类线索融合的优势,其改进效果有限。本文提出了一种新方法,学习将编码为体积表示的多视角与单目线索进行融合,无需启发式构建的掩码。我们的分析揭示,多视角线索在静态区域能捕捉更精确的几何信息,而单目线索在动态区域能捕捉更有用的上下文。为了让从静态区域多视角线索中学习到的几何感知传播到动态区域的单目表示中,并让单目线索增强多视角代价体积的表示,我们提出了一种跨线索融合(CCF)模块,其中包括跨线索注意力(CCA)机制,用于从每个来源编码空间非局部的内部相互关系,以增强另一来源的表示。在真实世界数据集上的实验证明了所提方法的显著有效性和泛化能力。