Remotely captured images possess an immense scale and object appearance variability due to the complex scene. It becomes challenging to capture the underlying attributes in the global and local context for their segmentation. Existing networks struggle to capture the inherent features due to the cluttered background. To address these issues, we propose a remote sensing image segmentation network, RemoteNet, for semantic segmentation of remote sensing images. We capture the global and local features by leveraging the benefits of the transformer and convolution mechanisms. RemoteNet is an encoder-decoder design that uses multi-scale features. We construct an attention map module to generate channel-wise attention scores for fusing these features. We construct a global-local transformer block (GLTB) in the decoder network to support learning robust representations during a decoding phase. Further, we designed a feature refinement module to refine the fused output of the shallow stage encoder feature and the deepest GLTB feature of the decoder. Experimental findings on the two public datasets show the effectiveness of the proposed RemoteNet.
翻译:遥感图像因复杂场景而具有极大尺度和目标外观变异性,这给在全局与局部上下文中捕捉其分割的潜在属性带来了挑战。现有网络受杂乱背景影响难以提取固有特征。为解决这些问题,我们提出了一种用于遥感图像语义分割的网络——RemoteNet。通过结合Transformer与卷积机制,我们提取了全局与局部特征。RemoteNet采用编码器-解码器架构并利用多尺度特征,构建了注意力图模块以生成通道注意力权重用于特征融合。在解码网络中设计了全局-局部Transformer模块(GLTB),以支持解码阶段的鲁棒表示学习。此外,我们提出特征精炼模块,用于优化浅层编码器特征与解码器最深GLTB特征的融合输出。在两个公开数据集上的实验结果表明了所提RemoteNet的有效性。