Semantic scene completion (SSC) jointly predicts the semantics and geometry of the entire 3D scene, which plays an essential role in 3D scene understanding for autonomous driving systems. SSC has achieved rapid progress with the help of semantic context in segmentation. However, how to effectively exploit the relationships between the semantic context in semantic segmentation and geometric structure in scene completion remains under exploration. In this paper, we propose to solve outdoor SSC from the perspective of representation separation and BEV fusion. Specifically, we present the network, named SSC-RS, which uses separate branches with deep supervision to explicitly disentangle the learning procedure of the semantic and geometric representations. And a BEV fusion network equipped with the proposed Adaptive Representation Fusion (ARF) module is presented to aggregate the multi-scale features effectively and efficiently. Due to the low computational burden and powerful representation ability, our model has good generality while running in real-time. Extensive experiments on SemanticKITTI demonstrate our SSC-RS achieves state-of-the-art performance.
翻译:语义场景补全(SSC)通过联合预测整个三维场景的语义与几何结构,在自动驾驶系统的三维场景理解中发挥着关键作用。得益于语义分割中的上下文信息,SSC已取得快速发展,但如何有效利用语义分割中的语义上下文与场景补全中的几何结构之间的关系仍待深入探索。本文从表征分离与BEV融合的角度出发,提出户外SSC解决方案。具体而言,我们提出SSC-RS网络,该网络采用具有深度监督的独立分支,显式解耦语义表征与几何表征的学习过程。同时,我们设计了配备自适应表征融合(ARF)模块的BEV融合网络,用于高效聚合多尺度特征。由于计算负担低且表征能力强,该模型在实时运行的同时展现出良好的泛化性。在SemanticKITTI数据集上的大量实验表明,SSC-RS达到了最先进的性能水平。