Event-based semantic segmentation has gained popularity due to its capability to deal with scenarios under high-speed motion and extreme lighting conditions, which cannot be addressed by conventional RGB cameras. Since it is hard to annotate event data, previous approaches rely on event-to-image reconstruction to obtain pseudo labels for training. However, this will inevitably introduce noise, and learning from noisy pseudo labels, especially when generated from a single source, may reinforce the errors. This drawback is also called confirmation bias in pseudo-labeling. In this paper, we propose a novel hybrid pseudo-labeling framework for unsupervised event-based semantic segmentation, HPL-ESS, to alleviate the influence of noisy pseudo labels. In particular, we first employ a plain unsupervised domain adaptation framework as our baseline, which can generate a set of pseudo labels through self-training. Then, we incorporate offline event-to-image reconstruction into the framework, and obtain another set of pseudo labels by predicting segmentation maps on the reconstructed images. A noisy label learning strategy is designed to mix the two sets of pseudo labels and enhance the quality. Moreover, we propose a soft prototypical alignment module to further improve the consistency of target domain features. Extensive experiments show that our proposed method outperforms existing state-of-the-art methods by a large margin on the DSEC-Semantic dataset (+5.88% accuracy, +10.32% mIoU), which even surpasses several supervised methods.
翻译:事件驱动语义分割因其能够处理传统RGB相机无法应对的高速运动与极端光照场景而受到广泛关注。由于事件数据难以标注,现有方法通常依赖事件到图像的重建来获取伪标签进行训练。然而,这不可避免地引入噪声,而基于噪声伪标签(尤其是单源生成的伪标签)的学习可能强化误差——这一缺陷在伪标签学习中被称为确认偏差。本文提出一种新颖的混合伪标签框架HPL-ESS,用于无监督事件驱动语义分割,以减轻噪声伪标签的影响。具体而言,我们首先采用朴素的无监督域适应框架作为基线,通过自训练生成一组伪标签;随后将离线事件到图像重建融入框架,通过预测重建图像的分割图获得另一组伪标签。为混合两组伪标签并提升质量,我们设计了一种噪声标签学习策略。此外,我们提出软原型对齐模块以进一步改善目标域特征的一致性。大量实验表明,本方法在DSEC-Semantic数据集上以显著优势超越现有最佳方法(准确率提升5.88%,平均交并比提升10.32%),甚至优于部分监督方法。