In computer vision, depth estimation is crucial for domains like robotics, autonomous vehicles, augmented reality, and virtual reality. Integrating semantics with depth enhances scene understanding through reciprocal information sharing. However, the scarcity of semantic information in datasets poses challenges. Existing convolutional approaches with limited local receptive fields hinder the full utilization of the symbiotic potential between depth and semantics. This paper introduces a dataset-invariant semi-supervised strategy to address the scarcity of semantic information. It proposes the Depth Semantics Symbiosis module, leveraging the Symbiotic Transformer for achieving comprehensive mutual awareness by information exchange within both local and global contexts. Additionally, a novel augmentation, NearFarMix is introduced to combat overfitting and compensate both depth-semantic tasks by strategically merging regions from two images, generating diverse and structurally consistent samples with enhanced control. Extensive experiments on NYU-Depth-V2 and KITTI datasets demonstrate the superiority of our proposed techniques in indoor and outdoor environments.
翻译:在计算机视觉领域,深度估计对于机器人、自动驾驶、增强现实和虚拟现实等方向至关重要。语义信息与深度的结合可通过信息双向共享来增强场景理解能力。然而,数据集中语义信息的匮乏带来了显著挑战。现有卷积方法受限于局部感受野,难以充分挖掘深度与语义之间的共生潜力。本文提出一种数据集无关的半监督策略以应对语义信息稀缺问题。该方法设计了深度语义共生模块,通过局部与全局上下文中的信息交换实现全面相互感知。此外,创新性地提出NearFarMix增强方法,通过策略性融合两幅图像的区域来生成结构一致且多样性增强的样本,从而缓解过拟合并补偿深度-语义任务。在NYU-Depth-V2和KITTI数据集上的大量实验表明,所提技术在室内外环境中均展现出优越性能。