Depth-aware panoptic segmentation is an emerging topic in computer vision which combines semantic and geometric understanding for more robust scene interpretation. Recent works pursue unified frameworks to tackle this challenge but mostly still treat it as two individual learning tasks, which limits their potential for exploring cross-domain information. We propose a deeply unified framework for depth-aware panoptic segmentation, which performs joint segmentation and depth estimation both in a per-segment manner with identical object queries. To narrow the gap between the two tasks, we further design a geometric query enhancement method, which is able to integrate scene geometry into object queries using latent representations. In addition, we propose a bi-directional guidance learning approach to facilitate cross-task feature learning by taking advantage of their mutual relations. Our method sets the new state of the art for depth-aware panoptic segmentation on both Cityscapes-DVPS and SemKITTI-DVPS datasets. Moreover, our guidance learning approach is shown to deliver performance improvement even under incomplete supervision labels.
翻译:深度感知全景分割是计算机视觉中的一个新兴课题,它结合语义与几何理解以实现更鲁棒的场景解析。近期研究致力于构建统一框架以应对这一挑战,但多数方法仍将其视为两个独立的学习任务,限制了跨领域信息探索的潜力。我们提出一种深度统一的深度感知全景分割框架,该框架以相同的目标查询,在逐片段方式下联合执行分割与深度估计。为缩小两个任务之间的差距,我们进一步设计了一种几何查询增强方法,能够利用潜在表示将场景几何信息融入目标查询。此外,我们提出一种双向引导学习方法,通过利用任务间的相互关联促进跨任务特征学习。我们的方法在Cityscapes-DVPS和SemKITTI-DVPS数据集上均实现了深度感知全景分割的最新最优结果。实验表明,即使在标注不完整的情况下,我们的引导学习方法仍能带来性能提升。