In recent years, the research community has shown a lot of interest to panoramic images that offer a 360-degree directional perspective. Multiple data modalities can be fed, and complimentary characteristics can be utilized for more robust and rich scene interpretation based on semantic segmentation, to fully realize the potential. Existing research, however, mostly concentrated on pinhole RGB-X semantic segmentation. In this study, we propose a transformer-based cross-modal fusion architecture to bridge the gap between multi-modal fusion and omnidirectional scene perception. We employ distortion-aware modules to address extreme object deformations and panorama distortions that result from equirectangular representation. Additionally, we conduct cross-modal interactions for feature rectification and information exchange before merging the features in order to communicate long-range contexts for bi-modal and tri-modal feature streams. In thorough tests using combinations of four different modality types in three indoor panoramic-view datasets, our technique achieved state-of-the-art mIoU performance: 60.60% on Stanford2D3DS (RGB-HHA), 71.97% Structured3D (RGB-D-N), and 35.92% Matterport3D (RGB-D). We plan to release all codes and trained models soon.
翻译:近年来,研究界对提供360度方向视角的全景图像表现出浓厚兴趣。通过输入多种数据模态并利用其互补特性,可基于语义分割实现更鲁棒、更丰富的场景理解,以充分挖掘其潜力。然而,现有研究主要集中于针孔RGB-X语义分割。在本研究中,我们提出一种基于Transformer的跨模态融合架构,旨在弥合多模态融合与全向场景感知之间的鸿沟。我们采用感知畸变模块来处理等矩形表示所导致的极端物体变形与全景畸变。此外,在特征融合之前,我们通过跨模态交互实现特征校正与信息交换,从而为双模态与三模态特征流传递长程上下文信息。在三个室内全景视角数据集中,利用四种不同模态类型的组合进行的全面测试表明,我们的技术取得了最先进的mIoU性能:在Stanford2D3DS(RGB-HHA)上达60.60%,在Structured3D(RGB-D-N)上达71.97%,在Matterport3D(RGB-D)上达35.92%。我们计划近期开源所有代码及训练模型。