Existing approaches focus on using class-level features to improve semantic segmentation performance. How to characterize the relationships of intra-class pixels and inter-class pixels is the key to extract the discriminative representative class-level features. In this paper, we introduce for the first time to describe intra-class variations by multiple distributions. Then, multiple distributions representation learning(\textbf{MDRL}) is proposed to augment the pixel representations for semantic segmentation. Meanwhile, we design a class multiple distributions consistency strategy to construct discriminative multiple distribution representations of embedded pixels. Moreover, we put forward a multiple distribution semantic aggregation module to aggregate multiple distributions of the corresponding class to enhance pixel semantic information. Our approach can be seamlessly integrated into popular segmentation frameworks FCN/PSPNet/CCNet and achieve 5.61\%/1.75\%/0.75\% mIoU improvements on ADE20K. Extensive experiments on the Cityscapes, ADE20K datasets have proved that our method can bring significant performance improvement.
翻译:现有方法侧重于利用类别级特征来提升语义分割性能。如何刻画类内像素与类间像素之间的关系,是提取具有判别性的代表性类别级特征的关键。本文首次提出通过多重分布来描述类内变化。随后,提出了多重分布表征学习(MDRL)方法,用于增强语义分割中的像素表征。同时,我们设计了一种类别多重分布一致性策略,以构建嵌入像素的判别性多重分布表征。此外,我们提出了一个多重分布语义聚合模块,用于聚合对应类别的多重分布,以增强像素语义信息。我们的方法可以无缝集成到流行的分割框架FCN/PSPNet/CCNet中,并在ADE20K数据集上实现了5.61%/1.75%/0.75%的mIoU提升。在Cityscapes、ADE20K数据集上的大量实验证明,我们的方法能够带来显著的性能提升。