Endeavors have been recently made to transfer knowledge from the labeled pinhole image domain to the unlabeled panoramic image domain via Unsupervised Domain Adaptation (UDA). The aim is to tackle the domain gaps caused by the style disparities and distortion problem from the non-uniformly distributed pixels of equirectangular projection (ERP). Previous works typically focus on transferring knowledge based on geometric priors with specially designed multi-branch network architectures. As a result, considerable computational costs are induced, and meanwhile, their generalization abilities are profoundly hindered by the variation of distortion among pixels. In this paper, we find that the pixels' neighborhood regions of the ERP indeed introduce less distortion. Intuitively, we propose a novel UDA framework that can effectively address the distortion problems for panoramic semantic segmentation. In comparison, our method is simpler, easier to implement, and more computationally efficient. Specifically, we propose distortion-aware attention (DA) capturing the neighboring pixel distribution without using any geometric constraints. Moreover, we propose a class-wise feature aggregation (CFA) module to iteratively update the feature representations with a memory bank. As such, the feature similarity between two domains can be consistently optimized. Extensive experiments show that our method achieves new state-of-the-art performance while remarkably reducing 80% parameters.
翻译:最近的研究致力于通过无监督域适应(UDA)将标注的针孔图像域知识迁移至未标注的全景图像域,旨在解决由等距柱状投影(ERP)中像素非均匀分布导致的风格差异和畸变问题造成的域间隙。以往工作通常基于几何先验知识,采用专门设计的多分支网络架构进行知识迁移,这导致大幅增加计算成本,且其泛化能力因像素间畸变差异而严重受限。本文发现ERP中像素的邻域区域确实引入更少畸变。据此,我们直觉性地提出一种新型UDA框架,可有效解决全景语义分割中的畸变问题。相比之下,我们的方法更简单、更易实现且计算效率更高。具体而言,我们提出无需几何约束即可捕获邻域像素分布的畸变感知注意力(DA)机制。此外,我们提出类别级特征聚合(CFA)模块,通过内存库迭代更新特征表示。由此,两个域间的特征相似性可得到持续优化。大量实验表明,本方法在减少80%参数量的同时实现了新的最优性能。