With the increasing availability of depth sensors, multimodal frameworks that combine color information with depth data are gaining interest. However, ground truth data for semantic segmentation is burdensome to provide, thus making domain adaptation a significant research area. Yet most domain adaptation methods are not able to effectively handle multimodal data. Specifically, we address the challenging source-free domain adaptation setting where the adaptation is performed without reusing source data. We propose MISFIT: MultImodal Source-Free Information fusion Transformer, a depth-aware framework which injects depth data into a segmentation module based on vision transformers at multiple stages, namely at the input, feature and output levels. Color and depth style transfer helps early-stage domain alignment while re-wiring self-attention between modalities creates mixed features, allowing the extraction of better semantic content. Furthermore, a depth-based entropy minimization strategy is also proposed to adaptively weight regions at different distances. Our framework, which is also the first approach using RGB-D vision transformers for source-free semantic segmentation, shows noticeable performance improvements with respect to standard strategies.
翻译:随着深度传感器的日益普及,结合颜色信息与深度数据的多模态框架正受到广泛关注。然而,语义分割的真实标注数据获取成本高昂,使得域适应成为一个重要的研究领域。但大多数域适应方法无法有效处理多模态数据。具体而言,我们针对无源域适应这一具有挑战性的场景展开研究——即在不重复使用源数据的情况下完成域适应。我们提出MISFIT(多模态无源信息融合Transformer),这是一种深度感知框架,通过在多阶段(即输入层、特征层和输出层)向基于视觉Transformer的分割模块注入深度数据。颜色与深度风格迁移有助于早期域对齐,而跨模态自注意力机制的重新连接可生成混合特征,从而提取更丰富的语义内容。此外,我们还提出了一种基于深度熵的最小化策略,用于自适应地对不同距离的区域进行加权。本框架作为首个将RGB-D视觉Transformer用于无源语义分割的方法,在性能上较标准策略表现出显著提升。