In this paper, we focus on semantic segmentation method for point clouds of urban scenes. Our fundamental concept revolves around the collaborative utilization of diverse scene representations to benefit from different context information and network architectures. To this end, the proposed network architecture, called APNet, is split into two branches: a point cloud branch and an aerial image branch which input is generated from a point cloud. To leverage the different properties of each branch, we employ a geometry-aware fusion module that is learned to combine the results of each branch. Additional separate losses for each branch avoid that one branch dominates the results, ensure the best performance for each branch individually and explicitly define the input domain of the fusion network assuring it only performs data fusion. Our experiments demonstrate that the fusion output consistently outperforms the individual network branches and that APNet achieves state-of-the-art performance of 65.2 mIoU on the SensatUrban dataset. Upon acceptance, the source code will be made accessible.
翻译:本文聚焦于城市场景点云的语义分割方法。我们的核心思想在于协同利用多种场景表征,以从不同的上下文信息和网络架构中获益。基于此,提出的网络架构APNet被分为两个分支:点云分支和航拍图像分支(其输入由点云生成)。为充分发挥各分支的不同特性,我们采用几何感知融合模块,通过学习将各分支结果进行融合。各分支独立的额外损失函数不仅避免了某一分支主导结果,还确保每个分支独立实现最优性能,并明确定义了融合网络的输入域,使其仅执行数据融合。实验表明,融合输出结果始终优于各独立网络分支,且APNet在SensatUrban数据集上实现了65.2 mIoU的当前最优性能。论文被接收后,源代码将公开提供。